2026-07-27
RAG optimization in practice: 6 key stages that determine retrieval accuracy
This article breaks down the 6 key stages in RAG optimization that determine retrieval accuracy: document parsing, chunking strategy, embedding model selection, hybrid retrieval, Rerank-based reranking, and context assembly. Each stage comes with ready-to-use parameter combinations and engineering implementation approaches, plus a troubleshooting SOP and an iteration priority path, to help engineers quickly locate retrieval bottlenecks and systematically improve RAG pipeline performance.
Stage 1: document parsing — the hidden culprit behind retrieval failures
When RAG results fall short, most teams' first instinct is to swap the embedding model or tune retrieval parameters. But if you have done systematic bad case attribution, you will find a less intuitive fact: a large share of retrieval failures originate at the ingestion stage, not the retrieval stage. Typical failure modes include PDF tables parsed into scrambled text, OCR output from scanned documents riddled with garbled characters, and Markdown heading hierarchies lost during conversion, breaking the context. No matter how sophisticated the downstream retrieval pipeline is, these problems cannot be remedied: the candidate pool simply contains no chunks of acceptable quality.
A practical troubleshooting rule: when the correct answer cannot be found in the Top-50 coarse retrieval results, don't rush to add Rerank or tune the embedding model. First go back and check whether document parsing and chunking have destroyed key information. Rerank can only reorder chunks that have already entered the candidate pool; it cannot rescue an answer that was lost at the source.
Parsing strategies for three document types
Different document types need different processing pipelines. Blindly running every format through the same flow is a common engineering mistake:
| Document type | Core risk | Recommended handling | Quality gate |
|---|---|---|---|
| Plain text / well-structured Markdown | Low risk; heading hierarchy occasionally lost | Ingest directly, preserving the heading hierarchy as metadata | Spot-check heading tree integrity |
| PDF/Word with tables, lists, and code blocks | Table rows and columns scrambled; list items merged into a single paragraph | Use a structured parser to extract boundaries and output type-annotated JSON | Manually spot-check table fidelity |
| Scanned / image-based PDF | OCR errors; formulas and special symbols garbled | Filter by character confidence after OCR engine processing | Mark paragraphs below the confidence threshold as "pending manual review" instead of ingesting them directly |
Engineering implementation essentials
The goal of structured parsing is not to "convert PDFs into text" but to "convert documents into structured data with boundary annotations." Specifically:
- Keep tables as standalone chunks: Preserve each table as an independent chunk, with its header information attached as metadata. Once a table is broken up and mixed into body text, its semantic integrity cannot be recovered.
- Keep lists atomic: If each item in an ordered list contains an independent condition or conclusion, keep the items within the same chunk so that conditions and conclusions are not split into different chunks.
- Tag code blocks by type: Code segments should be annotated with their language and ingested as complete units, excluded from the subsequent text chunking logic.
- Metadata travels with the chunk: Each output chunk carries fields such as source file name, page number/section path, document type, and parsing confidence, for filtering and source tracing during retrieval.
In the open-source toolchain, both Unstructured and MinerU can output structured results annotated with element types. When evaluating them, focus on two things: first, table reconstruction accuracy on the document formats most common in your business; second, whether they support custom post-processing pipelines (for example, routing low-confidence paragraphs to a manual review queue).
Deployment recommendations
Early in the project, spend one to two days on an "ingestion quality audit": randomly sample 50 document chunks, manually compare the source with the ingested text, and calculate the information loss rate. The ROI of this exercise is far higher than repeatedly tuning parameters at the retrieval layer. If the information loss rate exceeds 10%, every subsequent optimization is built on a flawed foundation. Fix the foundation first, then worry about the superstructure.
Stage 2: chunking strategy — parameter combinations you can copy directly
Chunking is the most underestimated stage in a RAG pipeline. Many teams drop documents in, slice them at a default fixed length with chunk_size set to 1000 and overlap set to 0, then spend enormous amounts of time tuning prompts, which gets things backward. Chunk quality directly sets the ceiling on retrieval precision; the embedding model and Rerank downstream can only optimize within that ceiling.
Why fixed-size chunking falls short
The core problem with fixed-length chunking is that it is completely blind to text structure. A complete set of operating steps may be split down the middle, and the first and second halves of a code block may land in different chunks, so a user's question can only retrieve fragments. Industry evaluations generally show that this naive approach achieves retrieval accuracy of roughly 60%, meaning nearly half of queries fail to get a useful answer. And the precision loss caused by improper chunk granularity (too fine or too coarse) can reach 30% to 50%, which is not something fine-tuning can recover.
Semantic-aware chunking: cut along structural boundaries
The fix is straightforward: make split points fall on the text's natural boundaries, such as the end of a paragraph, a heading change, or the close of a code block. This semantic-aware strategy does not require complex models; rules alone capture most of the gain. Reproductions across multiple open-source evaluations show that this adjustment alone can raise recall by about 15%–20%.
Going further, you can use sentence-level embedding similarity to detect whether the topic has shifted:
- Compute cosine similarity between adjacent sentences
- Similarity < 0.5 → treat as a topic shift and split here
- Similarity ≥ 0.5 → treat as the same topic and keep merging
The 0.5 threshold is an empirical starting point for Chinese text. English text has stronger semantic continuity between sentences, so the threshold can be raised as appropriate. In production, run a threshold sweep against labeled bad cases with a step size of 0.05; it usually converges within two or three rounds.
Recommended parameters: a Parent-Child two-tier structure
Single-tier chunks always face a tension: finer granularity makes retrieval precise but fragments the context sent to the LLM, while coarser granularity keeps context intact but adds retrieval noise. The Parent-Child strategy decouples this tension:
| Tier | Token length | Overlap | Purpose |
|---|---|---|---|
| Child chunk (for retrieval) | 300–400 tokens | 50–100 characters | Vectorization and similarity matching |
| Parent chunk (for context) | 1,200–2,000 tokens | About 200 characters | Once a child chunk is hit, send its parent chunk to the LLM |
In implementation, each child chunk records its parent chunk ID when stored. At retrieval time, child chunks are used for top-K retrieval; after deduplicating parent chunks, the full parent passages are assembled into the prompt. This ensures both precision at the retrieval stage (fine-grained matching) and context integrity at the generation stage (coarse-grained input).
Overlap prevents information loss at chunk boundaries. An overlap window of 50–100 characters is enough to cover a complete short sentence, preventing key information from being split across the edges of two chunks so that neither side can retrieve it.
Metadata enrichment: give each chunk a "retrieval wrapper"
Once the raw text is chunked, there is one more low-cost, high-return step: generate a concise summary title (kept under 50 characters) plus keywords for each chunk, and concatenate them into the retrieval text for vectorization.
Why does this work? User questions are usually highly abstract (such as "how do I configure an SSL certificate"), while the raw chunk may consist entirely of concrete steps and never contain the words "configure SSL certificate." The summary title effectively gives the chunk an additional semantic entry point in the retrieval space. Industry testing generally shows that this enrichment brings a significant further gain in accuracy on top of semantic chunking.
Implementation tip: use a lightweight model (or rule-based extraction of the top 10 multi-character terms by TF-IDF) to generate keywords, and use an LLM to generate titles in a single batch. Prepend the title and keywords to the chunk body, separated by a line break, and use the whole as the embedding input. Note that this enrichment text is used only for vectorization, not in the context ultimately sent to the LLM, which avoids introducing noise.
Summary of the decision path
- Step 1: Replace fixed-size chunking with rule-based chunking along paragraph, heading, and code block boundaries to capture the baseline gain at zero cost
- Step 2: Add sentence embedding similarity for topic boundary detection, starting with a threshold of 0.5
- Step 3: Adopt the Parent-Child two-tier structure, using 300–400-token child chunks for retrieval and 1,200–2,000-token parent chunks for generation
- Step 4: Append a summary title and keywords to each chunk to widen retrieval coverage
These four steps build on one another, and each can be validated independently. After each step, run your evaluation set and confirm that recall is actually rising before moving on. This avoids over-engineering.
Stage 3: embedding model selection — decision tree and parameter cheat sheet
The embedding model is the most expensive component to replace in the retrieval pipeline. Switching models means re-vectorizing every document, rebuilding the index, and cutting live traffic over through a phased rollout. By comparison, adjusting chunking parameters only requires rerunning the back half of the pipeline. So the selection decision needs to be made correctly early in the project, rather than relying on "let's try a different model" as a late-stage fallback.
Selection decision tree
Narrow your scenario down to one of three paths and evaluate them in turn:
| Scenario | Recommended model | Output dimensions | Max input tokens | Core trade-off |
|---|---|---|---|---|
| Primarily Chinese, with chunks kept within 512 tokens | Chinese-specific embedding model | — | — | Low cost for on-premises deployment; mature Chinese semantic representation |
| Mixed Chinese-English corpus, or very long passages that need a large window | Multilingual embedding model | — | — | Produces both dense and sparse representations, a natural fit for hybrid retrieval |
| Accuracy first; API latency and cost are acceptable | Commercial high-dimensional embedding model | — | — | Strong discrimination in high-dimensional space; supports on-demand dimensionality reduction to compress storage |
Order of evaluation: first confirm language distribution and document length, then look at deployment constraints (whether external APIs can be called and whether GPU resources are available). Most internal enterprise knowledge bases are primarily in Chinese and do not exceed 512 tokens after chunking, so the first path covers them.
Key parameter cheat sheet
- Dimensions: There are three tiers: 768 / 1024 / 3072. Higher dimensions give stronger fine-grained semantic discrimination, but index size and retrieval latency grow linearly. In real projects, 1024 dimensions is the price-performance inflection point; any marginal gain beyond that needs to be validated on your own evaluation set.
- Max input length: Models with a 512-token limit require upstream chunking to control length strictly. Anything beyond the limit is truncated rather than raising an error, which is a hidden recall loss. If your chunking strategy runs long (800+ tokens), you must choose a model with an 8192-token window.
- Normalization: Most models output L2-normalized vectors by default, in which case cosine similarity is equivalent to the inner product and Inner Product can be used for retrieval. Note that some models (such as early versions of M3E) require a manual normalization step; otherwise, similarity scores are not comparable.
- Quantization: FP16 is the deployment baseline. Compressing further to INT8 halves the index size, and industry tests generally find the accuracy loss negligible. But the benefit of quantization depends on data distribution, so compare Recall on your own evaluation set before deciding whether to ship it.
Selection principle: evaluate first; switching models is the last resort
The right workflow is:
- Run Recall@50 with the current model and current chunking (take the Top-50 candidates and check whether the target document is among them) to establish a baseline.
- If Recall@50 is already high enough but final answers are inaccurate, the bottleneck is most likely in Rerank or context assembly, not the embedding model.
- Only consider switching models when Recall@50 is clearly insufficient and there is still no improvement after adjusting the chunking strategy and adding metadata filtering.
The hidden costs of switching models are easy to underestimate: the compute cost of re-embedding every document, doubled storage while old and new indexes coexist, and coordinated changes to dimension parameters across upstream and downstream pipelines. For a document library in the millions, a single model switch typically takes weeks of engineering time. Compare that cost with the cost of "trying a few more sets of chunk parameters," and the priority becomes clear.
A practical rule of thumb: if your evaluation set isn't built yet, don't agonize over model selection. Start with BGE-large-zh or BGE-M3 and spend your effort building 50–100 labeled Query-Document pairs. With an evaluation set, every decision can be backed by data; without one, any model switch is blind tuning.
Stage 4: hybrid retrieval — engineering dual-path retrieval with vectors + BM25
Pure vector retrieval has a structural blind spot: it scores by semantic similarity, but for exact identifiers such as error codes (e.g., ERR_0x80070005), SKU numbers, and software version numbers, embedding models tend to map them to nearby regions of the vector space, so results get mixed up. The engineering answer is not a bigger model but an additional BM25 keyword retrieval path as an exact-match fallback.
Standard dual-path retrieval flow
The engineering skeleton is intuitive:
- Vector retrieval takes the Top-50 candidates
- BM25 keyword retrieval takes the Top-50 candidates
- Merge and deduplicate the two result sets, capping the candidate pool at about 100
- Fuse and rank the merged results with RRF (Reciprocal Rank Fusion)
- Pass the fused, ordered list to the downstream Rerank module for fine ranking
RRF's calculation logic is simple: for each document, take its rank in each result list and sum 1/(k + rank) as the final score. The parameter k controls how smoothly the rank-based decay is applied. In practice, most teams start tuning from a middle value, which is also the default in several mainstream frameworks. Within a reasonable range, k does not affect the final ranking dramatically, so there is no need to spend much effort tuning it early on; focusing on the quality of each retrieval path yields bigger returns.
Query Rewrite: always keep the original query
Query Rewrite is a common technique for improving vague user questions, but it introduces an easily overlooked risk: the rewrite model itself may misread the user's intent. Once a rewrite goes off course, the entire downstream pipeline (retrieval, ranking, generation) drifts with it, and the drift is subtle because the system "appears to be running normally" while simply answering the wrong question.
The engineering defense is simple: always retrieve with both the original query and the rewritten query, then fuse the two result sets. That way, even if the rewrite is wrong, results from the original query remain in the candidate pool as a fallback. The implementation cost is almost zero, just one extra retrieval request, yet it effectively prevents a rewrite error from collapsing the entire pipeline.
Multi-Query: split one question into multiple retrieval directions
When a user asks a compound question, a single query often hits only some of the relevant documents. The idea behind Multi-Query is to have the LLM expand one question into 3–5 sub-queries from different angles, retrieve for each one, and then merge and deduplicate the results.
For example, if a user asks "What exactly is your SLA?", the question can be expanded into:
- The committed service availability percentage
- Fault response and recovery time requirements
- Breach compensation clauses
- Specific definitions of each service tier
Each of the four retrieval paths takes its own Top-N; after merging and deduplication, the candidate pool covers far more ground than a single query. This is especially effective when information in the knowledge base is scattered across multiple documents.
Engineering considerations for deployment
| Concern | Recommendation |
|---|---|
| BM25 index maintenance | Update it in sync with the vector index; when documents change, rebuild both paths to avoid data inconsistency |
| Deduplication strategy | Deduplicate by chunk_id; different chunks of the same document count as different candidates |
| Number of Multi-Query sub-queries | 3–5 is appropriate; more increases retrieval latency with diminishing returns |
| Query Rewrite model choice | You don't need the strongest model; a lightweight model (e.g., GPT-3.5 class) is enough. The key is keeping the original query as a fallback |
| Latency control | Retrieval paths can be issued in parallel, so total latency depends on the slowest path rather than the sum of all paths |
The core value of hybrid retrieval is not that it "adds one more path to pure vector search," but that it uses engineering to hedge against the systemic weaknesses of any single retrieval paradigm. Semantic understanding and exact matching are complementary capabilities, and in enterprise knowledge base scenarios there is almost never a case that needs only one of them. Building both retrieval paths solidly and giving the downstream Rerank a high-quality candidate pool is the prerequisite for the entire RAG pipeline to perform well.
Stage 5: reranking — the critical hop from coarse filtering to fine ranking
Vector retrieval is essentially a two-tower architecture: the query and the chunk are each encoded independently into vectors, and cosine similarity is then calculated. The inherent flaw of this architecture is that fine-grained interaction information between the two texts is lost at the encoding stage. For example, if the query asks "How long does the confidentiality obligation survive after the contract is terminated?", vector retrieval can pull back every passage containing "confidentiality" and "contract termination," but vector distance alone cannot reliably rank which one actually answers the core intent of "survival period."
A Cross-Encoder solves exactly this problem: it concatenates the (Query, Chunk) pair into a single sequence and feeds it to a Transformer, letting the model fully cross-compute the semantic relationship between the two texts in every attention layer and output a fine-grained relevance score. The trade-off is that nothing can be precomputed. Every pair must pass through the model, so it can only be used after the candidate pool has already been narrowed.
Tiered Top-K: a three-level funnel from coarse retrieval to fine ranking to the prompt
The most common engineering mistake is setting only one Top-K parameter. A sound approach is to split the retrieval chain into three levels:
| Stage | Number of candidates | Role |
|---|---|---|
| Coarse retrieval (after vector + BM25 fusion) | 30–100 chunks | Ensure recall by getting the correct answer into the pool |
| Kept after Rerank | 5–10 chunks | Reorder by semantic relevance and cut noise |
| Final input to the prompt | 3–6 chunks | Control token cost and reduce the chance of the model being distracted by irrelevant passages |
The benefit of this funnel is twofold: accuracy at the top positions improves significantly after fine ranking, and because only a few high-quality chunks are fed into the prompt at the end, the number of tokens sent to the generation model drops sharply, reducing both latency and cost.
When to use Rerank, and when not to
The criterion is simple: first check whether the correct answer is present at the coarse retrieval stage (Top-50):
- The correct chunk is in the candidate pool but ranked low (for example, at positions 15–40): This is the classic use case for Rerank. The problem is insufficient ranking precision, and a Cross-Encoder can move the chunk up.
- The correct chunk is not in the candidate pool at all: Rerank cannot conjure an answer out of nothing. Go back and investigate upstream: whether document parsing lost content, whether chunking severed key information, whether metadata filtering is too aggressive, and whether query rewriting drifted from the original intent. Until those stages are fixed, adding Rerank just means carefully sorting a garbage heap.
A simple diagnostic: run Recall@50 on the evaluation set. If it is already above 90% but Precision@5 is unsatisfactory, add Rerank directly; if Recall@50 itself is below 70%, investigate the retrieval pipeline first.
Engineering implementation essentials
- Model choice: For Chinese, bge-reranker-v2-m3 and BAAI/bge-reranker-large are both proven options; for English, the ms-marco family of Cross-Encoders remains a stable baseline.
- Trading off candidate pool size against latency: Cross-Encoder inference time scales linearly with the number of candidates. With the candidate pool kept under 50 and GPU inference, the latency of a single Rerank pass stays within an acceptable engineering range; above 100 candidates, run a lightweight pre-filter first before sending them to fine ranking.
- Score threshold vs. fixed count: In practice, combine the two: take the Top-N first (e.g., 5), then set a minimum score threshold to filter out clearly irrelevant results. This avoids situations where "the last two of the Top-5 have extremely low scores but still get stuffed into the prompt."
Rerank is one of the highest-ROI single-point optimizations in the entire retrieval pipeline: no index rebuild and no change to the chunking strategy, just one extra fine-ranking step at inference time, and the quality of the context that ultimately enters generation goes up a clear notch. But it presupposes that coarse retrieval has already caught the correct answer. If the first layer of the funnel leaks, no finer sieve downstream will help.
Stage 6: context assembly and generation constraints — last-mile engineering details
Even if the first five stages have retrieval and ranking well tuned, problems in the context ultimately assembled for the model will erode that earlier work. This section covers the easily overlooked engineering details between "retrieval results" and "the prompt": what to pass, how much to pass, and how to keep the model from making things up.
Parent-child chunk retrieval: retrieve precisely, feed the model completely
Anyone who has tuned chunking has run into this tension: small chunks give precise retrieval hits but fragmented content, so the model sees only half a sentence; large chunks keep semantics intact but get diluted by irrelevant information during retrieval, lowering the hit rate. There are two common approaches to context enrichment in the industry. One is sentence window expansion: after a sentence is hit, expand it by a few sentences on either side before passing it to the model. The other is parent-child chunk (parent document) retrieval. Across comprehensive evaluations, the parent-child approach usually performs better on both retrieval precision and context integrity, making it the one worth deploying first.
The approach is to build a two-level index: at chunking time, use small chunks (for example, a few hundred characters each) to build the vector index, and match against small chunks during retrieval. But once a small chunk is hit, what the system actually returns to the LLM is not the small chunk itself but the larger parent chunk that contains it (possibly an entire document or chapter). For example, a 3,000-character product manual is split into six 500-character child chunks, each indexed separately; when a user's question hits the third child chunk, what is ultimately passed to the model is the full text of the manual, not that isolated 500-character fragment. The retrieval stage enjoys the precise matching of small chunks while the generation stage loses no context, so you get the best of both. In engineering terms, you need to maintain the child-to-parent mapping and handle the retrieval and generation layers separately. The implementation cost is modest, which makes this a high-value optimization.
Context compression: trade one extra LLM call for cleaner context
Parent-child chunks address "whether to pass complete passages"; context compression addresses "how much of what gets passed in is noise." Concretely, before generating the final answer, have an LLM (or a lighter model) read through the retrieved chunks, extract the sentences that are truly relevant to the user's question, discard irrelevant background and repetitive statements, and then assemble the condensed content into the final prompt.
This method significantly reduces the share of noise in the context, with especially noticeable results when chunks are large and contain a lot of unrelated content. The cost is one extra LLM call, which means extra latency and cost. So this stage should not be enabled indiscriminately. It fits best where answer accuracy is critical and users can accept waiting an extra second or two, such as legal clause verification and contract review. For consumer-facing instant Q&A, the latency of an extra call may not be worth it; you can trigger it on demand only after parent-child retrieval rather than turning it on by default across the entire pipeline.
Two constraints at the generation stage: low temperature + citation format
Once the context is assembled, what remains is constraining generation itself. Both of these measures are cheap and pay off quickly, so there is almost no reason not to adopt them.
- Keep temperature in the low range of 0.1–0.3. A RAG scenario calls for "faithfully relaying the retrieved information," not free improvisation by the model. Lowering the temperature noticeably reduces random drift in answers, so asking the same question several times yields more consistent answers, which matters in production environments that need reproducibility and debugging.
- Explicitly constrain the citation format in the prompt, for example by requiring the model to mark, for every conclusion, which retrieved chunk and which document it came from. This is not about looks but about troubleshooting: once an answer is wrong, you can pinpoint directly whether the retrieval stage retrieved the wrong content or the generation stage misread correct retrieval results, and the two problems call for completely different fixes. Without this constraint, when something goes wrong you can only comb through retrieval logs by hand one entry at a time, which is very inefficient.
The combined value of these steps
Viewed individually, none of these stages (parent-child retrieval, context compression, low temperature settings, citation constraints) delivers a dramatic improvement on its own. But feedback from real deployments shows that when semantic chunking, hybrid retrieval, reranking, context compression, and prompt constraints are chained together, the gains compound: recall can rise from below 50% to near-perfect, accuracy can rise in step from just over 60% to above 90%, response time can be cut by nearly half, and some token consumption can be saved as well. This is why we don't recommend optimizing only one or two stages: single-point optimization easily hits a bottleneck, and only tuning the pipeline as a whole produces a step change in results.
A word of caution: more stages are not always better. Compression and Rerank both add latency, and how far to go depends on how the business weighs accuracy against response time. The troubleshooting SOP in the next section gives more specific guidance on priorities.
Deployment path: troubleshooting SOP and iteration priorities
Optimizing a RAG system is not a one-time tuning exercise but an engineering process of converging stage by stage. The most common failure mode in practice is skipping foundational data governance and piling on Rerank and prompt tricks, only to find after a great deal of effort that the problem was key passages lost at the parsing layer. Below is a proven deployment sequence and troubleshooting path.
Recommended iteration order
The steps are ordered from highest to lowest ROI. At each step, quantify the change in Recall@K and end-to-end accuracy, and confirm the gain before moving on to the next:
| Phase | Action | Core goal |
|---|---|---|
| 1 | Data governance | Ensure documents are fully ingested, with no lost passages or garbled text, and with tables and images handled |
| 2 | Build a minimum evaluation set | Establish a quantifiable baseline and end subjective "it feels better" judgments |
| 3 | Compare chunking strategies | Run A/B tests with the evaluation set to find the best chunking parameters for the current corpus |
| 4 | Hybrid Search | Vector + BM25 dual paths complement each other and close the blind spots of single-path retrieval |
| 5 | Query Rewrite | Handle query quality issues such as colloquial phrasing, unclear references, and multiple intents |
| 6 | Reranking | Push the correct evidence to the top positions, provided it is already in the candidate pool |
| 7 | Context compression | Send fewer noisy chunks to the generation model, lowering hallucination risk and token consumption |
| 8 | Generation constraints | Prompt-level safeguards: citation marking, refusal rules, and format control |
| 9 | Phased rollout monitoring | Continuously collect bad cases in production and loop back to phase 2 |
The logic of this order: the first four steps address "can the correct answer get into the candidate pool," and the last five address "once it is in the pool, can it be used correctly." Reversing the order leads to a lot of wasted tuning.
Minimum evaluation set: how to build it
A starting size of 50–100 items is enough. The key is to cover four types of scenarios:
- High-frequency questions: the top queries in production logs, representing your core traffic
- Known failure cases: questions users have reported as "wrong answer" or "not found," representing current weaknesses
- Exact-match questions: queries involving error codes, version numbers, SKU numbers, and other items that must be correct character for character
- Should-refuse questions: questions for which the knowledge base truly has no answer, to verify whether the system hallucinates a response
Each evaluation item contains three fields: question (the user's original question), golden_context (the source document chunk expected to be hit), and golden_answer (the reference answer). golden_context lets you measure retrieval-layer recall on its own, rather than diagnosing retrieval and generation problems together.
Bad case troubleshooting decision tree
When an evaluation item fails, locate the bottleneck along the following path:
- Step 1: Check whether the coarse retrieval candidate pool (Top-50) contains the correct evidence
- If it doesn't, the problem is upstream in the retrieval layer. Check in order: was the document ingested → did parsing drop passages → did chunking break up the answer → is metadata filtering too strict → does the query need rewriting → is a BM25 channel needed
- If it does, the problem is in the ranking or generation layer. Check in order: did Rerank rank the correct chunk too low → was the context sent to the LLM truncated → does the prompt lack citation constraints or refusal rules
The core principle of this rule: when the candidate pool does not contain the answer, never reach for Rerank first. Rerank can only reorder existing candidates; it cannot retrieve missing content out of thin air. Focus on getting the correct evidence into the pool first, then think about how to rank it higher.
Measurement discipline for step-by-step convergence
After completing each phase of changes, rerun the evaluation set and record two core metrics: Recall@5 (whether the Top-5 contains golden_context) and end-to-end accuracy (whether the final answer matches golden_answer). Keep a change only when it brings an observable metric improvement; otherwise, roll it back. This "change one step, measure one step" discipline avoids the dilemma of being unable to attribute results when multiple variables change at once.
FAQ
What is the right chunk_size?
There is no universally optimal value, but there is an actionable framework for choosing one. Start with the characteristics of your corpus. Technical documentation usually has longer paragraphs and strong context dependencies, so larger chunks (in the 500–700 token range) preserve semantic integrity. FAQs and customer service conversations are inherently short texts, so smaller chunks (200–350 tokens) better match their native granularity. Long-paragraph texts such as legal contracts can go above 800 tokens, because breaking them up would damage the integrity of clauses. Once you have a candidate range, the definitive answer comes from comparison experiments on your evaluation set. Don't trust any "universal best value."
If I already use vector retrieval, do I still need BM25?
Almost always. Vector retrieval excels at semantic matching (synonyms, or asking the same thing in different words), but it is unstable in exact keyword scenarios: when a user searches for an error code, product model, or proper noun, the embedding model may map it to a chunk that is semantically similar but factually different. BM25's term-frequency matching fills exactly this gap. In engineering terms, running two independent retrieval paths and fusing them with RRF is cheap, while the recall gain for exact-match queries is often substantial. If your evaluation set contains "error code/version number" questions and single-path vector retrieval performs poorly, adding BM25 is essentially a guaranteed win.
Will Rerank slow down response times?
It depends on candidate pool size and model choice. A Cross-Encoder runs one forward pass for each (query, chunk) pair; with the candidate pool kept at 20–50, the latency that mainstream Rerank models add to overall response time is acceptable. If you are extremely latency-sensitive (for example, a conversational scenario that requires P99 under 2 seconds), you can shrink the candidate pool to 20 or choose a lightweight Rerank model. Another engineering trick is asynchronous prefetching: start coarse retrieval while the user is still typing, so the perceived wait is shorter by the time Rerank runs. In short, the accuracy gain from Rerank is usually worth the small latency cost. The real question is whether the candidate pool contains the correct answer, not how fast the ranking is.
How do I decide which stage of my current system to optimize first?
Go back to the bad case troubleshooting decision tree. Take the 20–30 most recent failure cases and label each one with the stage its failure belongs to (parsing loss, over-fragmented chunks, retrieval miss, low ranking, generation hallucination, failure to refuse). Whichever stage has the highest failure count is the one to tackle first. This is far more reliable than choosing a direction by intuition. And if you don't have an evaluation set yet, your first priority is building the evaluation set itself: without measurement, there is no optimization.