Teverant AI · Insights

2026-06-08

Enterprise RAG pitfalls: why knowledge base projects fail and how to fix them

Why do enterprise knowledge base RAG projects demo well and then fail as soon as they go live? This article systematically examines six core pitfalls (document parsing, vector retrieval performance, low recall, data governance decay, missing permission boundaries, and the absence of an evaluation framework) and, drawing on real cases, lays out actionable countermeasures such as hybrid retrieval, access control, and quantitative evaluation, helping technical teams avoid detours and move enterprise RAG systems from demo to production.

Why your RAG project dies after the demo: a map of failure modes

We have seen too many projects like this: the demo runs so smoothly that the business side signs off on the spot, and three months later the system is quietly taken offline, with "inconsistent results" given as the reason. Stability problems can never be summed up in a single sentence, but the paths to failure are strikingly similar.

Looking back at these projects, almost every team made the same conceptual mistake: they treated the RAG project as an evaluation of large language model (LLM) capabilities rather than as a data engineering project. As a result, resources went almost entirely into model selection and prompt tuning, while the data pipeline was cobbled together with scripts, document parsing ran on default parameters, and the retrieval architecture was copied from tutorial examples. During the demo phase all of these problems stayed hidden, because the POC documents were hand-picked: clean formatting, focused content, a manageable volume.

The landmines planted during the POC

A typical failure timeline looks like this: the POC runs on 80 to 100 documents, retrieval is accurate, the answers are usable, and the demo passes on the first try. At launch, the real document repository is imported—100,000 documents in mixed formats, including scanned files, nested tables, historical versions, and duplicate content—and recall falls off a cliff. Users complain that "nothing it answers hits the point," the operations team starts manually maintaining "answers to frequently asked questions," and the knowledge base degrades into a static FAQ page with a search box.

The problems do not arise after launch; they are already lurking during the POC. Curated documents sidestep parsing challenges, a small vector database masks retrieval performance problems, and uniform formatting works around flaws in the chunking strategy. Once the real environment is introduced, every hidden weakness is exposed at the same time.

The structural causes of failure

Break down a large number of failed cases and the problems cluster along three links in the chain:

  • The data pipeline is not robust. The reality of enterprise documents is this: PDFs contain scanned images, Word files contain floating tables, key information in PowerPoint decks is hidden inside graphics, and wiki pages are nested with external links. When this content is processed directly with a general-purpose parsing library, the output text is fragmented and semantically broken; once it enters the vector database, it is effectively noise rather than knowledge. Parsing quality sets the ceiling on knowledge, yet in the vast majority of projects this step amounts to a single line of code: call some open-source library, pass in the file path, take the output.
  • The retrieval architecture breaks down at scale. Pure vector similarity retrieval performs well on small-scale queries with clear semantics. But a substantial share of real enterprise user questions contain exact product model numbers, clause numbers, personal names, and dates—information that vector retrieval almost never hits reliably, so keyword retrieval must be added as a complement. Once scale goes up, retrieval latency, index update frequency, and the reranking weights for multi-path results each become an engineering decision that needs its own design, not something default parameters can handle.
  • Context governance is missing. A knowledge base is not a static asset. Policies are revised, products iterate, old documents are never taken down, new documents carry no expiration date—and the answers users receive may be based entirely on obsolete information. This problem is not obvious early after launch; it starts to erupt three to six months in, surfacing as business feedback that "the answers don't match reality," while the technical team has no mechanism at all to pinpoint which document is contaminating the results.

The counterintuitive reality of resource allocation

Industry surveys consistently show that in RAG systems that have actually made it to production at scale, the largest contributors to final quality are data quality and retrieval architecture, while the model itself contributes relatively little. This ratio is the exact inverse of how most teams actually allocate resources—typically more than 70% of the effort goes into model evaluation and prompt engineering, while data and retrieval are handled on a "get it running first, worry later" basis.

This is not to say the model doesn't matter; rather, the model's ceiling is determined by retrieval quality. Feed a top-tier model context that has been misparsed and semantically fragmented, and its output quality will not repair itself just because it has a large parameter count—it will very fluently package the noise into wrong answers that look plausible. That is more dangerous than an outright error, because users won't notice right away, and trust erodes quietly.

Working back from failure modes to engineering priorities

At project kickoff, several questions deserve clear answers before the first line of code is written:

  • What is the format distribution of the real document repository? What share are scanned files? Are there nested tables or graphical content? This determines the complexity budget for the parsing solution.
  • What proportion of user queries require exact matching (model numbers, clauses, dates)? This determines whether hybrid retrieval must be designed in from day one.
  • How often are documents updated, and what is the expiration mechanism? Who decides when a document should be removed from the knowledge base? This determines the design of the data governance process.
  • Are there permission requirements—are some documents retrievable only by specific roles? If so, permission filtering at the retrieval layer must be built in during architecture design; retrofitting it later is practically infeasible.

There are no universal answers to these questions, but if all of them are skipped during the POC, every one will turn into a separate incident ticket after launch. The sections that follow break down the specific engineering decisions for each stage, along with acceptance metrics that can be quantified.

Document parsing: the first overlooked pitfall

When many teams review a failed RAG project, the problems point to the embedding model, prompt design, or retrieval algorithm—but few look further upstream. Document parsing and chunking are the foundation of the entire pipeline, and the quality of that foundation sets the ceiling for every optimization that follows. No amount of work on recall can make up for semantics that were already broken at ingestion.

The corpus reality: messy formats are the norm, not the exception

The corpus of a real enterprise knowledge base differs enormously from a lab test set. Industry surveys consistently show that unstructured documents make up the bulk of enterprise knowledge bases: PDF contracts, Excel reports, Markdown technical docs, scanned files, text exported from PPT... the number of formats easily exceeds twenty. This means you cannot assume "clean plain-text input"—every format has its own parsing traps: multi-column PDF layouts interleave the left and right columns into scrambled sentences; Excel cell relationships lose their context once extracted as text; and scanned files run through OCR introduce character-level noise.

Errors at the parsing stage have one especially dangerous property: they fail silently. Documents are ingested normally, vectors are generated normally, and retrieval returns results—except those results contain sentences in scrambled order, truncated table rows, or terms that OCR misread. Problems like these are hard to catch in a demo, because demos tend to use the most neatly formatted documents.

The cost of hard splits: 32% of semantics truncated before ingestion

After parsing comes chunking. The most common shortcut is to split by a fixed token count—say, a cut every 512 tokens. The problem with this strategy is not that it is "fixed" as such, but that it is completely blind to text structure: a complete line of argument, a numbered procedural step, or a table's title and its data rows can all be cut in half at a split point.

According to industry test data (no specific report is cited, so this is downgraded to a qualitative statement), the rate of semantic truncation under fixed-size chunking is quite high and seriously undermines subsequent retrieval accuracy. The one figure with a clear source comes from a technical comparison test: after an adaptive chunking algorithm was introduced, semantic integrity improved by 65% compared with fixed-size chunking, with chunk lengths concentrated in the 300–800 token range (data source: an industry RAG evaluation; the specific report name will be added once verified).

The core idea of adaptive chunking is to use semantic boundaries, not character counts, as the basis for splitting. Common boundary signals in practice include section headings, blank lines, the end of a list item, and a period followed by a line break. The splitting logic itself is not complicated; the hard part is that documents in different formats require different boundary-detection rules, which need to be customized for your corpus.

Engineering rules of thumb for chunk parameters

There is a relatively stable industry consensus on chunking parameters:

  • Chunk length: 300–800 characters (tokens). Chunks under 300 characters often lack sufficient context, and their vector representations tend to drift; beyond 800 characters, the core semantics are diluted and matching precision at retrieval time drops.
  • Overlap: 10%–20%. Keeping some overlap between adjacent chunks prevents a split point from severing a logical relationship that spans sentences. Too little overlap loses information at the boundaries; too much causes duplicate content to recur in retrieval results and interfere with ranking.
  • Splitting priority: semantic boundaries > sentence boundaries > length limits. The length limit is a fallback constraint, not the primary basis for splitting.

Engineering check: use a distribution chart to surface hidden problems

There is a simple way to quickly verify the quality of document parsing and chunking: compute the length distribution of every chunk ingested and plot it as a histogram. Under a sound strategy, the distribution should concentrate in the target range with a relatively tight shape. If you see either of the following two anomalies, something in the pipeline is wrong:

  • Large numbers of very short chunks (< 100 characters): usually the result of parsing failures (such as misordered multi-column PDFs or table cells split row by row) or overly aggressive chunking logic. Once ingested, these chunks not only have little retrieval value but also dilute the quality of retrieval results.
  • Large numbers of oversized chunks (> 1000 characters): a sign that semantic boundary detection is not working and large sections of documents are being stuffed wholesale into a single chunk. The vector representation of an oversized chunk is a blend of several semantic topics, and its matching precision at retrieval time is very poor.

This check costs almost nothing. Run it once before the knowledge base goes live and it will intercept most parsing-level problems. Rather than repeatedly tuning parameters after recall falls short, doing acceptance on ingestion quality up front is the more cost-effective engineering choice.

Only once parsing and chunking are done well does what enters the vector database become a truly meaningful semantic unit, and only then do downstream retrieval optimization, hybrid retrieval, and reranking have room to work. Skipping this step and going straight to tuning the model is like building on sand.

Vector retrieval performance traps: everything surfaces once you scale

Many teams are quite satisfied with retrieval speed during the POC—results in a few hundred milliseconds, a smooth demo. But after launch, as document volume grows to the 100,000 level, response times suddenly degrade to seconds or longer. This is not an occasional environment issue; it is debt that the architecture choices ran up, coming due once they meet scale.

An order-of-magnitude misperception

When estimating the load on a vector database, engineers habitually think in units of "documents," and that is the first cognitive trap. After chunking, 100,000 documents typically produce somewhere between 500,000 and 3 million chunks actually written to the vector database—the exact multiple depends on the chunking strategy and average document length. In other words, you think you are managing 100,000 records, but you are actually querying a vector collection on the order of millions. This gap is hidden at small scale and exposed the moment scale arrives.

The fundamental problem with flat indexes at scale

A flat index works by brute-force traversal: every query computes the distance between the query vector and every vector in the database. With an embedding dimension of 1024, the computation for a single query grows linearly with the number of chunks. When three conditions hold at once—millions of chunks, 1024 dimensions, and no sharding—slowness is a physical constraint, not a configuration problem, and no amount of parameter tuning will save you.

The more hidden cost is this: even if your hardware keeps average latency in check, P99 latency will still deteriorate sharply as soon as concurrency rises a little. A single user in a demo won't reveal it; multi-user concurrency in production exposes it immediately.

Index structure is the decisive variable

The empirical rule is that at a scale of more than 100,000 documents, the choice of index structure determines roughly 80% of retrieval performance, while all other factors combined (hardware, network, caching) account for the remaining 20%. This means the marginal return on hardware investment is far lower than that of switching the index structure from flat to a type suited to your scale.

There are two mainstream directions:

  • HNSW (Hierarchical Navigable Small World): graph-based approximate nearest neighbor search that reduces query complexity from linear to logarithmic, suited to scenarios that demand high recall accuracy and relatively ample memory. The trade-off is heavy memory usage during index construction, and it does not handle frequent incremental writes well.
  • IVF-PQ (Inverted File + Product Quantization): first uses an inverted file to partition the vector space into clusters, so that queries search only the relevant clusters; PQ then further compresses the vectors through product quantization. Suited to very large-scale, memory-constrained scenarios, it trades a small amount of precision for large gains in speed and storage.

Which one to choose depends on your SLA and resource constraints, but both have a fundamental advantage over flat indexes at scale. Before settling on a solution, first measure your total chunk count and target P95 latency, then work backward to the index type and shard count—rather than picking first and seeing how it performs.

Choosing dimensions: higher is not always better

Embedding dimensionality directly affects storage, compute, and memory bandwidth, but the gains from higher dimensions show clear diminishing returns. A/B test data from a financial company (source: an industry case study; the specific report name has not been disclosed) shows that raising dimensionality from 512 to 768 improved recall by about 12% but added 45ms to P95 latency. Pushing further to 1024 narrows the recall gain even more, while the latency cost keeps accumulating linearly.

Quantization is another technique worth prioritizing. In the same set of tests, applying quantization reduced model size by 70% with only a 3% loss in retrieval precision. For most enterprise knowledge base scenarios this trade-off pays off—a 3% precision gap is nearly imperceptible to the business, while a 70% reduction in size directly relieves memory and disk pressure.

Database partitioning: an underrated engineering lever

Mixing documents from every line of business into a single vector database is another common mistake. The problem has two layers: first, noise at query time—finance documents and research and development documents occupy very different semantic spaces, and mixing them lowers retrieval relevance; second, operations become too coarse-grained, so a surge in data volume in any one business domain affects performance across the board.

Partitioning by business domain or document type keeps each sub-database within a manageable size limit and lowers the cost of building and rebuilding indexes. Partitioning does not make query logic more complex—deciding which database to route to based on metadata at the routing layer is far easier than cleaning up a chaotic single database after the fact.

Pre-launch checklist

CheckTargetCommon omission
Load-test P95/P99 latencyMeet the SLA at target concurrencyTesting only average latency and ignoring the tail
Index rebuild planRebuild periodically to prevent fragmentationBuilt but never maintained; performance degrades within six months
Chunk count estimateEstimate the actual write volume based on the chunking strategyUsing document count instead of chunk count for capacity planning
Dimension and quantization validationMeasure recall and latency on the target datasetAdopting the model's default dimensions without benchmarking
Partition boundariesEach business domain or document type stands aloneThe POC's single database carried over into production

Index fragmentation is a chronic problem: frequent incremental writes gradually degrade the HNSW graph structure, and IVF cluster centroids drift as well. Without a periodic rebuild mechanism, 6 months later you will find that the same query is 30% slower than at launch, and no one can say when it started. Making index rebuilds part of routine operations is no more optional than VACUUM is for a database.

Four real causes of low recall, and a hybrid retrieval solution

The most common complaint after a RAG project goes live is: "The knowledge is clearly in there, but the system just can't find it." This problem usually doesn't surface in the demo—there are few documents and the questions are controlled. Once in production, gaps in recall are amplified by scale. In practice, troubleshooting shows that low recall concentrates in four places, and the fix for each is entirely different.

Four root causes

  • The chunking strategy destroys semantic integrity. The most common approach is to split text into blocks by a fixed character count, with the result that a complete explanation gets cut in half: the first half lands in chunk 37 and the second half in chunk 38, and neither is sufficient when retrieved on its own. Correct chunking needs to be aware of document structure—the boundaries of headings, paragraphs, and lists are the semantic boundaries, and character count is only a supporting constraint, not the primary basis.
  • The embedding model doesn't match the business corpus. General-purpose embedding models perform well on general semantics, but enterprise knowledge bases are full of industry jargon, internal code names, and product abbreviations. In a general model's vector space, these terms are either scattered or sit right next to unrelated concepts. The fix is domain fine-tuning on the business corpus, or at least choosing a model pretrained in a similar domain, rather than applying one trained for open-domain question answering as-is.
  • Queries and documents are not in the same semantic space. Users tend to ask conversational questions, whereas knowledge base documents are written as formal declarative statements. When both are encoded with the same embedding model, vector distance does not necessarily reflect semantic relevance. The problem is especially pronounced in asymmetric retrieval scenarios. Rewriting and expanding the query on the query side (query expansion), or using a bi-encoder trained specifically on question-answer pairs, are relatively direct remedies.
  • Vector retrieval only, with no keyword fallback. Vector retrieval excels at capturing semantic similarity but is inherently "face-blind" to exact strings—error codes (such as ERR_CONNECTION_RESET), contract numbers, product model numbers, and specific figures cannot be located reliably in vector space. When a user searches for "fault code E0047," vector retrieval may return a pile of content related to "troubleshooting" but not that exact record. This is not a question of model quality; it is a ceiling set by how vector retrieval works.

Hybrid retrieval: dual-path retrieval plus funnel reranking

For the fourth problem above, BM25 sparse retrieval is a necessary complement, not an optional optimization. It works on term-frequency statistics and is inherently sensitive to exact string matches, filling precisely the blind spot of vector retrieval. When both paths retrieve in parallel and the results are merged and deduplicated, coverage is significantly higher than with any single retrieval method.

But even as retrieval coverage increases, the context window sent to the LLM remains limited; you can't stuff dozens of candidates into it. This calls for a second-stage reranking step to do the compression. The recommended architecture is as follows:

StageMethodOutput sizeGoal
Dual-path retrievalVector retrieval + BM25 in parallel; results merged and deduplicatedTop 50–100Maximize coverage, reduce missed retrievals
Reranking (Rerank)A Cross-Encoder scores each query–candidate chunk pairTop 5–10Compress noise, increase relevance density
GenerationInsert the reranked results into the prompt sent to the LLMFinal answerReduce hallucination, improve accuracy

The essential difference between a Cross-Encoder and a Bi-Encoder is this: a Bi-Encoder encodes the query and the document separately and then compares their distance—fast, but with a ceiling on precision; a Cross-Encoder concatenates the query and document and runs them through a full attention computation—markedly more precise but computationally heavy, and unsuitable for searching the entire database directly. Placing the Cross-Encoder in the second stage of the funnel, where the candidate set has already been narrowed to a few dozen, is a sensible division of labor that balances precision and latency.

Reranking is, without exception, the single change that delivers the most visible improvement in enterprise RAG engineering. Industry observations from AI customer service deployments in finance show that after dual-path retrieval and reranking were introduced, end-to-end answer accuracy rose by about 30% and the fully automated resolution rate across the whole pipeline climbed to 75%. A jump of that size means the system has moved from "good enough for a demo" to genuinely production-grade.

Engineering decision checklist

  • When chunking, rely primarily on document structure boundaries and secondarily on character limits; keep a sliding-window overlap for long paragraphs to avoid truncating key information across chunks.
  • When evaluating embedding models, run a retrieval benchmark on your own business corpus—don't just look at public leaderboard rankings.
  • On the query side, consider query expansion or HyDE (Hypothetical Document Embeddings) to narrow the gap in semantic expression between queries and documents.
  • BM25 retrieval is standard equipment; especially when the knowledge base contains many model numbers, codes, and numeric values, leaving out BM25 will inevitably cause exact-match queries to fail.
  • The choice of rerank model and the size of the retrieval window (the K in Top K) need to be tuned together: if K is too small, the reranker has too few candidates; if K is too large, latency rises. 50–100 is a reasonable starting point for most scenarios.
  • After launch, continuously monitor recall@K at the retrieval layer rather than looking only at user ratings of the final answers—problems in the generation layer and in the retrieval layer look very similar on the surface, but the fixes point in completely different directions.

Data governance: keeping the knowledge base from rotting the moment it launches

One type of incident is especially awkward in a post-mortem: the system is running normally, the retrieval logic is fine, the embedding model hasn't degraded—yet the business side is already on fire. A retail company's case is a classic lesson: its customer service knowledge base went live without an update mechanism ever being put in place. When the platform later changed its refund timeframe rules, the old answers kept being retrieved in production, and in the several business days after the new rules took effect, daily complaints exceeded 300. The root cause was not the model; it was an outdated document.

This case supports a clear judgment: data decay is not technical debt; it is business risk. Technical debt can be scheduled for repayment; business risk incurs losses in real time. The handling priorities of the two differ by an order of magnitude.

Decay is outpacing the refresh capacity of traditional ETL

The traditional model for maintaining a knowledge base is the periodic full rebuild: export the documents, chunk them, re-embed, and overwrite the vector database. Where knowledge changes infrequently, this process is barely adequate. But the pace of business has changed. Industry surveys consistently show that enterprises' tolerance for knowledge update cycles has shrunk from quarterly maintenance in the past to a requirement that changes be reflected in near real time. A quarterly ETL batch cadence is a structural mismatch for these needs, not something parameter tuning can fix.

Full rebuilds also carry a hidden cost: each run involves a complete embedding computation and index rewrite. Once the document count reaches the tens of thousands, a single rebuild takes hours, during which the knowledge base is partially unavailable or in a muddle of versions. Triggering full rebuilds frequently amounts to deliberately creating service outage windows.

CDC: turning incremental changes into incremental updates

The solution is to switch the knowledge update path from "batch full rebuild" to "event-driven incremental updates." Change data capture (CDC) is the core component of this path. CDC works by listening to the write-operation logs of source systems (a database's binlog, a document system's change event stream) and extracting each insert, update, or delete as an independent change event; downstream consumers process these events to make the corresponding localized updates to the vector database instead of triggering a full rebuild.

This places several specific requirements on the architecture:

  • Source systems must be able to emit a structured change stream. Databases usually already can (MySQL binlog, PostgreSQL logical replication); unstructured document systems need hooks added at the write layer.
  • The vector database needs to support single-record updates and deletes by document ID or chunk ID, not just append-only writes. Some vector databases from early deployments support only batch writes, which is one source of migration cost.
  • Change events need to carry enough context for downstream consumers to determine the scope of impact: which document changed, which field changed, and which business domain the new version belongs to.

The results are quantifiable. After systematically introducing CDC-driven incremental synchronization together with metadata governance, one company compressed its knowledge update cycle from 72 hours to 15 minutes and reduced the error rate of knowledge base answers from 18% to 1.2% (data source: an industry case study whose specific report name has not been disclosed; downgraded here to a case description). An update latency of 15 minutes already clears the "acceptable" threshold for most business scenarios.

Metadata tagging: making invalidation and rebuilds surgical

CDC solves the problem of "getting changes to the vector database," but how precisely they are handled once there depends on the quality of the metadata attached to each chunk. Here is a set of tagging standards that, from an engineering standpoint, must be completed at ingestion:

Metadata fieldPurposeConsequence if missing
Document version numberDetermine whether a chunk comes from the latest versionOld and new content coexist; retrieval results are inconsistent
Last-updated timestampSupport recency-based filtering and down-weightingExpired content cannot be excluded at the retrieval layer
Business domain tagLimit the scope of impact during localized invalidationA single field change triggers a rebuild across all domains
Source document IDPrecisely locate and delete all chunks of a parent documentOrphan chunks linger after the document is deleted

The value of metadata is most evident during invalidation. After the refund-timeframe rule is updated, the ideal process is: CDC captures the change → the business domain tag is used to locate all chunks in the "refund policy" domain → only this batch of chunks is re-embedded and overwritten → all other content is left untouched. The time and compute cost of this process is a small fraction—on the order of a few percent—of a full rebuild. Without metadata, this kind of surgical, localized rebuild is impossible, and you fall back to sledgehammer mode: change one line of a rule, rebuild the entire knowledge base.

Engineering checklist

  • Confirm whether source systems can already emit change events and assess the cost of CDC integration—prioritize this ahead of optimizing the knowledge base itself.
  • At the chunking stage, enforce four metadata fields—version number, timestamp, business domain, and parent document ID—and do not accept null values.
  • When selecting a vector database, verify support for single-record updates and field-filtered deletes; these are prerequisites for CDC-based incremental synchronization.
  • Set an update-latency SLA for each business domain and establish alerts: notify proactively when the change-event backlog exceeds a threshold, rather than waiting for complaints before investigating.
  • Conduct regular "knowledge freshness audits": sample the version distribution of retrieval results to identify expired chunks that are still being retrieved.

Engineering investment in data governance tends to reach the agenda later than model tuning and retrieval optimization, but in production its risks are the first to surface. From the first day a knowledge base goes live, the decay clock is already ticking.

Permission boundaries: the most easily forgotten, highest-risk engineering item

Permission problems in RAG projects follow a typical exposure path: during the demo, a single account runs against all the data and everything works fine; after launch, multiple roles and departments come on board, and a salesperson sees contract clauses they should never have seen, or internal compensation data turns up in an external partner's query results. Only then does the team realize that permissions were never designed as an engineering problem—they were just a few lines of conditional rendering at the presentation layer.

This is one of the costliest architectural misjudgments in RAG systems.

Presentation-layer filtering ≠ permission isolation

Many teams' first instinct is: after retrieval, just hide whatever shouldn't be shown based on the user's role. The fundamental problem with this approach is that it hides only the interface output; the unauthorized content has already been sent into the LLM's context window. The model has already "read" this content while generating its answer, and even if the final response doesn't quote it directly, information leakage has in effect already occurred at the semantic level.

The correct place for isolation is the retrieval layer, not the rendering layer. Permission checks must happen at the filtering stage of vector retrieval, ensuring that every text chunk entering the context is one the current user is authorized to access. This is not a performance optimization issue; it is a compliance red line.

Permission granularity must go down to the field level

In enterprise settings, the same document is often visible to different roles to different extents. For a procurement contract, legal needs to see the full terms and the liability for breach, the business owner needs only the delivery dates and amounts, and external auditors may be allowed to see only a redacted summary. If the vector database stores just a single access tag for the whole document, you either over-restrict (even legal can't see everything) or over-expose (the business side sees confidential clauses).

In practice, permission labeling needs to be done during chunking at the document parsing stage—each chunk carries its own permission metadata instead of inheriting a single document-level tag. The vector database's metadata fields store these tags, and at retrieval time they are enforced as a hard filter rather than factored in as a ranking weight (soft rerank). The difference: soft ranking only pushes unauthorized content lower in the ranking, where it can still appear in the top-k results; a hard filter excludes illegitimate candidates before any distance is computed.

Multi-tenant isolation is mandatory, not optional

For SaaS deployments, or when multiple subsidiaries within a group share one knowledge base, tenant isolation must be implemented at the vector storage layer. There are two common engineering approaches: one is to create a separate collection or index partition for each tenant—physical isolation, so queries never cross boundaries; the other is to use tenant_id as a metadata filter within the same collection—logical isolation, which costs less but requires rigorous verification that the filter cannot be bypassed.

Which approach to choose depends on the number of tenants and the sensitivity of the data. Financial, healthcare, and legal scenarios should usually opt for physical isolation, not only for security but also for traceability in compliance audits—when regulators ask you to prove that data has not leaked across organizations, physical partitions are more convincing than filter logs.

Compliance constraints can't rely on prompts alone

Compliance control at the generation layer is another frequently underestimated link. In heavily regulated industries, errors in RAG system output are not just a user experience problem but a direct compliance risk and a source of financial loss. Over 18 months of serving clients in finance, healthcare, and law, Teverant AI has observed that for RAG systems in critical domains, every 1% reduction in error rate avoids an average of roughly RMB 580,000 in potential losses—a magnitude that makes "we added a disclaimer to the prompt" an untenable risk management strategy.

Business rules need to be embedded in the generation pipeline as code, not as natural language. Typical practices include: structurally validating retrieved content before generation (confirming the timeliness and authority of cited sources), filtering output through a rules engine after generation (blocking specific sensitive statements or conclusions beyond the authorized scope), and routing high-risk output to a human review queue instead of returning it directly. Prompts can serve as a supplement, but they cannot be the only line of defense.

Engineering checklist

  • Map out a permission matrix for every document type, specifying which roles have read access to which fields, and complete chunk-level permission labeling at the parsing stage
  • Store permission tags in vector database metadata, require the retrieval interface to receive the user's identity context, and enforce filters as hard conditions rather than ranking factors
  • In multi-tenant scenarios, choose physical or logical isolation based on data sensitivity, and run regular penetration tests targeting filter bypass
  • Run regression tests in staging with cross-permission queries: use a low-privilege account to construct queries that should be blocked, and verify that the retrieval results contain no unauthorized chunks
  • Introduce a rules engine at the generation layer, moving compliance constraints out of prompts and into versioned, auditable code logic
  • Establish permission-related monitoring metrics—the number of unauthorized retrievals intercepted and the trigger rate of compliance filters at the generation layer—with alerts on abnormal fluctuations

Permission boundaries are the highest-risk engineering item in a RAG project because their failures are usually silent—no errors, no performance degradation, the system appears to be running normally—until a data breach exposes the problem. Moving permission design forward to the architecture review stage, rather than patching it after launch, is basic engineering discipline for projects like these.

Evaluation: without quantitative metrics, there is no direction for optimization

The most insidious failure mode of a RAG project is not a crash but "it feels okay"—the team tunes parameters on subjective impressions and claims progress with every change, yet can't say how much progress was made, where, or at what cost. This is essentially flying blind: no instrument panel, only the pilot's intuition.

Four layers of metrics: miss one and you have a blind spot

Evaluation of enterprise RAG must cover four layers, each corresponding to a different point of failure in the system:

  • Retrieval layer — Recall@K: the proportion of relevant documents retrieved among the top K chunks returned. If this number is low, every downstream step is building on a faulty foundation, and neither reranking nor generation can compensate.
  • Ranking layer — MRR / NDCG: where the relevant results rank. If recall is acceptable but ranking is poor, the model reads noisy chunks first and answer quality collapses just the same. MRR suits scenarios where the goal is to "find the first correct answer"; NDCG suits complex question answering that requires synthesizing multiple passages.
  • Answer accuracy: whether the final generated answer is correct. This can be labeled manually, or semi-automated rules such as "the answer contains the key facts" can be used for batch pre-screening, followed by manual spot checks. This layer maps directly to the quality users perceive.
  • Latency — P95 / P99: don't look only at average latency; the long tail is what enterprise users actually experience. If P95 exceeds 3 seconds, the workflow will be abandoned no matter how high the accuracy is.

These four layers are connected in series: missed retrievals at the retrieval layer drag down answer accuracy, and noise at the ranking layer drives up generation latency (the model has to process more irrelevant context). Any layer without numbers behind it means optimization of that stage has no stopping condition.

Why the old metrics fall short

The BLEU score measures surface word overlap and inherently penalizes cases where the same correct answer is expressed in different words. Subjective experience ratings depend on the evaluator's state and on which samples are chosen, and cannot support a fair comparison between two architecture iterations. Both methods have their value on academic benchmarks, but in enterprise settings they don't answer the questions the business actually cares about.

What the business really needs to know is: after using this knowledge base, can an employee who looks up a question make the right decision? For process steps that depend on the knowledge base, how much has the task completion rate improved? Does the same type of question get consistent answers at different points in time (long-term consistency)? These three dimensions—task completion rate, decision accuracy, and long-term consistency—are the numerator of the ROI calculation. Without these numbers, all you can tell the CIO is that "user feedback has been good," which makes a weak case for renewing the budget.

A minimum viable approach to establishing an evaluation baseline

There's no need to wait until the system is complete to start evaluating. The lowest-cost starting point is to build a "golden question set": select 200–500 questions with clear reference answers from real business scenarios, covering three categories—high-frequency queries, edge cases, and multi-hop reasoning. Reference answers are confirmed by domain experts and include both key fact points (for semi-automated scoring) and reference document IDs (for retrieval-layer scoring).

The core value of this question set is regression testing. Every time you adjust the chunking strategy, swap the embedding model, or change the reranker threshold, you must run it against the golden question set and output a before-and-after comparison across all four layers of metrics. An architecture change without comparison data amounts to experimenting in production, with real users bearing the cost.

A common hidden risk is the "single-metric optimization trap": cutting chunks finer can raise Recall@K, but it also scatters relevant documents across more chunks, lowering MRR and hurting answer coherence. Unless both layers of metrics are monitored at the same time, this cost is completely invisible. Regression testing on the golden question set exists precisely to catch these cases where "optimizing A quietly breaks B."

Putting it into practice: build evaluation into CI/CD

Evaluation should not be an ad hoc step before launch; it should be a mandatory gate on every code merge. Specifically:

  • Add the golden-question-set evaluation script to the CI pipeline so that it runs automatically every time a PR is merged into the main branch.
  • Set minimum thresholds for each layer of metrics (e.g., Recall@5 ≥ 0.75, P95 latency ≤ 2s); PRs that fall below a threshold may not be merged.
  • Every PR description for an architecture change must include a metrics comparison table, so that reviewers can see the gains and costs of the change directly rather than relying on the submitter's subjective account.
  • Periodically (monthly is recommended), expand the golden question set with newly collected real query logs, so that the evaluation set doesn't gradually drift away from the actual distribution of business questions.

Building an evaluation framework looks like extra overhead early in a project, but what it actually does is turn "tuning by feel" into "deciding by data." A RAG system without a quantitative baseline accumulates technical debt with every iteration—you don't know which change planted a hidden problem until one day a particular scenario blows up. Treating the evaluation pipeline as infrastructure rather than an option is one of the prerequisites for a RAG project to survive past the demo stage.

FAQ

The RAG POC worked well, so why did recall drop noticeably after launch?

This is the classic "demo trap," and the root cause almost always lies in two dimensions: data scale and data quality.

A POC typically ingests only a few dozen to a few hundred documents, and they are often hand-screened "clean documents." At that scale, the vector space has little noise and clear semantic boundaries, and just about any embedding model produces results that look good. Once document volume grows to tens or even hundreds of thousands after launch, the picture is completely different:

  • The vector space becomes denser: document fragments that are semantically similar but carry different answers proliferate, the Top-K results contain more distractors, and the accuracy of LLM-generated answers declines accordingly.
  • Document quality is not controlled: scanned files, badly formatted PDFs, and duplicate versions are ingested in bulk; the parsed chunks are noise to begin with, and even the most precise retrieval can't save them.
  • The chunking strategy isn't adjusted to content type: fixed-length chunking got by in the POC, but after launch, when it meets table-heavy financial reports or technical manuals with long procedures, fixed splits sever key context and semantic integrity is lost.

The remedy: first run a data audit to pick out poorly parsed documents for separate handling; then define chunking strategies by content type; finally, add hybrid retrieval (vector + BM25) and reranking at the retrieval layer to counter the noise caused by vector space density. Don't expect that switching to a better embedding model will solve it—the model is just one variable.

How should you choose an embedding model? Is the top of the leaderboard good enough?

Leaderboard scores are measured on general benchmark datasets, which often don't share the distribution of your business corpus. There are plenty of real-world cases where a team dropped in the leaderboard's number one and it fell flat.

Several practical constraints should factor into model selection:

  • Language and domain fit: vertical domains such as Chinese law, healthcare, and finance have large numbers of specialized terms, and general multilingual models often fail to represent these terms distinctly enough. If your documents are highly specialized, prioritize models fine-tuned on domain corpora, or be prepared to do the fine-tuning yourself.
  • Trade-off between vector dimensions and retrieval latency: higher-dimensional vectors are more precise for similarity computation, but index size and retrieval latency both increase. At scale, the latency difference between 1536 and 768 dimensions becomes very noticeable at p99.
  • Maximum token length: some models truncate chunks longer than 512 tokens, distorting the encoding of information at the end of long documents. Align chunk size with the model's effective context window.
  • Feasibility of private deployment: enterprise knowledge bases usually can't send document content to external APIs and need models that can run on the internal network. The closed-source model at the top of a cloud leaderboard is ruled out immediately here.

The safest approach: build a small evaluation set from your own business corpus (a few hundred question-answer pairs is enough), measure recall and MRR on the candidate models, and make decisions based on measured data rather than leaderboard screenshots.

At which layer should document permissions be enforced? Isn't front-end filtering enough?

Front-end filtering is presentation-layer control, not security control. The essential difference is this: the front end only takes care of "not displaying it to the user," but the data has already been pulled out by the retrieval layer—as soon as someone bypasses the front end and calls the API directly, or a bug appears in the front-end logic, the permission boundary is breached.

Proper access control should happen at the retrieval layer, and there are two mainstream implementation paths:

  • Metadata filtering: tag each document chunk's metadata with permission labels (department, role, classification level, etc.), and at retrieval time pass the permission conditions corresponding to the user's identity into the vector database as filter conditions, so that only chunks with matching permissions enter the retrieval candidate pool. This approach is widely supported by today's mainstream vector databases, and the performance overhead is manageable.
  • Multi-index isolation: build a separate vector index for each permission domain, with different users querying different indexes. Isolation is complete, but index maintenance costs are high and the storage overhead becomes significant at large document volumes; it suits scenarios with few permission domains and very clear boundaries.

One often-overlooked detail: when generating an answer, the LLM synthesizes multiple retrieved fragments. If the retrieval layer doesn't filter by permission, the LLM may blend content from documents the user isn't authorized to see into its answer—the final answer the user sees already contains information they shouldn't see, yet your logs show a "normal retrieval." This kind of leak is very hard to detect in an audit.

Bottom line: access control must be pushed down to the retrieval layer; front-end filtering can serve only as a UI convenience, never as the backbone of the security mechanism.

The knowledge base is updated frequently. Does the index need a full rebuild every time?

No, and it shouldn't. Full rebuilds are extremely costly at large document volumes and create windows during which the indexing service is unavailable. From an engineering standpoint, the system should be designed around incremental updates from the start.

Key points for incremental updates:

  • Unique document IDs and version tracking: tag each document at ingestion with a unique ID and a version number (or content hash). On update, compare hashes first; only documents whose content has actually changed trigger re-parsing and re-embedding.
  • Chunk-level granularity: if only page 3 of a 50-page manual has changed, ideally only the corresponding chunks are updated rather than the entire document being reprocessed. This requires establishing a mapping from pages or sections to chunks at the document parsing stage.
  • Deleting old vectors can't be overlooked: when a document is modified or retired, the chunk vectors of the old version must be removed from the index; otherwise the old content will keep being retrieved and produce wrong answers. Many projects handle only additions and forget the cleanup, and over time the index fills up with invalid content.
  • Asynchronous update queues: in high-frequency update scenarios, document change events go into a message queue and background workers handle parsing and embedding asynchronously, so that updates don't block the main flow or cause jitter in the retrieval service.

If incremental updates weren't considered when the system architecture was designed, retrofitting them later is quite expensive. This is a classic "neglect it early, pay for it later" engineering item. We recommend building document ID management and version tracking into the data model design at the very start of the project, rather than scrambling to fix things once full rebuilds have become unacceptably slow.