2026-09-22
What is RAG: how to implement enterprise knowledge base Q&A
What is RAG? Covering document processing, indexing and retrieval, query handling and reranking, answer generation, evaluation, and access governance, this article systematically explains how to implement question answering over an enterprise knowledge base and put it into production.
1. Should your enterprise adopt RAG? Diagnose the problem before choosing a vector database
To decide whether you need RAG, start by looking at how your knowledge changes, whether answers must be traceable, and whether different users should see the same content. Vector databases, embedding models, and rerankers are only implementation components; they can't substitute for judging the scenario. If the problem doesn't require retrieving external knowledge in the first place, building a complex index up front only adds more things to maintain.
RAG is better suited to factual knowledge that "keeps changing and must be backed by a source." Typical scenarios include questions about internal policies, product troubleshooting support, engineering documentation lookup, and verifying contract and compliance clauses. This content often never made it into a general-purpose model's training data, or it continues to be revised frequently after the model was trained. In these cases, the system should read the currently valid materials each time a request comes in and preserve the correspondence between the answer and the source text, rather than trying to make the model permanently memorize enterprise facts.
| Problem characteristics | Preferred approach | Engineering judgment |
|---|---|---|
| Very little material, processed once, no access control needed | Use a long context directly | A short pipeline; usually no need to build an index and update mechanism |
| Documents keep growing, questions recur | Standard RAG | Narrow the evidence first, then pass it to the model for generation, reducing irrelevant input |
| Need a fixed tone, structured format, or consistent execution of certain actions | Prompt constraints or fine-tuning | You're adjusting model behavior, not injecting frequently changing facts |
| Answers depend on cross-document relationships and multi-step connections | Evaluate GraphRAG | The cost of a graph structure is justified only when standard retrieval can't handle relational reasoning |
Long context and RAG are not substitutes for each other. For summarizing, comparing, or reviewing a few fixed documents, handing the full text directly to the model is usually simpler. But as the volume of material grows, stuffing everything into the context consumes more tokens and can dilute key information with large amounts of irrelevant content. If materials also need to be filtered by department, project, or classification level, the retrieval layer becomes a necessary control point. A common practice is to first retrieve a small amount of highly relevant evidence and then let the model perform the synthesis within a longer context.
Nor should fine-tuning be treated as a channel for updating enterprise knowledge. Once a policy version, pricing rule, or product specification changes, preparing new samples and retraining the model is slow, and it's hard to confirm that the old facts have been fully overwritten. Fine-tuning is better suited to standardizing wording, output fields, refusal behavior, and task workflows. A sounder division of labor: RAG supplies facts that can be updated and cited, while fine-tuning or prompt engineering constrains how the model uses those facts.
GraphRAG, in turn, addresses a different kind of problem: the answer doesn't exist directly in any single passage but requires linking up relationships among people, organizations, equipment, events, and so on scattered across documents. Related work from Microsoft Research has demonstrated an approach that uses entity-relationship graphs to support reasoning across materials. But graph extraction has to deal with entity disambiguation, relationship validation, and incremental maintenance, and the query pipeline is longer. Therefore, unless evaluation has already shown that vector retrieval has a consistent gap on multi-hop questions, GraphRAG should not be the default architecture.
Before kicking off a project, four questions make a quick screen: Is the knowledge updated frequently? Must answers cite their sources? Do materials need to be isolated according to user permissions? Does a single question involve only a local portion of the documents? When most of the answers are "yes," RAG is usually worth the investment; if you only occasionally process a small amount of fixed text, starting with long context is often more economical. Architecture choices should start from accountability for answers and the knowledge lifecycle, not from database types.
2. Document processing: most RAG problems arise before retrieval even starts
The offline processing in RAG can't be reduced to "extract the text, cut it into pieces, compute vectors." A retrieval system can only use information that has made it into the index: if heading relationships are broken, table fields are misaligned, and outdated policies are never taken down, then even if you later switch retrieval algorithms or large language models (LLMs), you'll only be searching for answers in a distorted corpus. From an engineering standpoint, establish a unified data ingestion pipeline first, and only then discuss index parameters.
The first step is to convert different sources into an intermediate format that is unified without losing structure. PDFs, Word documents, web pages, emails, tickets, and database records are stored differently; after parsing, you need at a minimum to distinguish body text, headings, lists, tables, attachments, and page positions. Scanned files need OCR first, and page numbers or region coordinates should be preserved to make it easy to locate the source text. Tables can't simply be strung together in visual order, or the model may be unable to tell which column a value belongs to; a safer approach is to combine the table name, column headers, row identifiers, and cell contents into records that can be understood on their own.
| Processing step | Common distortion | Recommended check |
|---|---|---|
| Format parsing | Headings demoted to body text; list order lost | Are hierarchy, numbering, and paragraph types preserved? |
| OCR | Model numbers, units, and figures misrecognized | Are low-confidence regions reviewed, and can the original image be traced? |
| Table conversion | Column headers separated from data rows | Can a single text chunk explain what its fields mean? |
| Version handling | Old and new policies both take part in retrieval | Are version relationships, effective periods, and repeal status clearly defined? |
The second step is cleaning, but the goal of cleaning is not to make the text "look tidy"; it's to remove content that would mislead retrieval. Headers, footers, copyright boilerplate, navigation menus, and email signatures may repeat on every page and, because they're so frequent, tend to crowd the retrieval results; garbled text and parsing fragments degrade the quality of keyword and vector representations; and duplicate files, expired pages, and invalidated clauses can lead the system to conclusions that no longer apply. Before deleting anything, keep the original files and processing records so that the data lineage can be reproduced if a dispute arises.
Each document and text chunk should also carry metadata that can be used to determine its scope of applicability, such as the source address, file identifier, version number, publication and effective dates, owning department, confidentiality level, applicable products, and parse time. Metadata is not an afterthought; it is the foundation for retrieval filtering, answer citations, incremental updates, and access control down the line. If only the body text is stored, the system will struggle to distinguish "similar in content" from "usable by the current user and still valid."
The third step is chunking according to content structure. A fixed character length can serve as a fallback, but it shouldn't be the uniform rule for every document. Policy documents should be split along chapters, clauses, and exception conditions, so that constraints aren't left behind in the previous chunk; operating manuals should keep action steps, prerequisites, risk warnings, and result checks within the same semantic unit as far as possible; tickets can be organized around the problem, the troubleshooting process, and the resolution; and tables need to keep column definitions together with the corresponding data. Chunk size also needs to fit the context limits of the embedding model and the generation model, rather than copying some generic parameter.
When chunks are too short, retrieval may hit a one-sentence conclusion while missing the subject it applies to, a negating condition, or the time range; when chunks are too long, large amounts of irrelevant text dilute the topic and take up more context during generation. To judge whether chunking is reasonable, spot-check samples: taken out of the original document, can any given chunk still answer "what is it describing, who does it apply to, and under what conditions"? If not, you need to merge in context or add structural information.
In production, you can store both fine-grained child chunks and more complete parent content. The retrieval stage uses smaller units to improve precision in locating content, and after a hit, the system fetches the chapter it belongs to, the adjacent steps, or the full clause and passes that to the model. This avoids the topic blurring that comes from matching whole chapters against vectors, and it reduces the risk of the model answering on the basis of an isolated sentence. Parent–child relationships, paragraph order, and source positions should be established at ingestion, not guessed at on the fly when generating answers.
The standard for document processing being complete is not "the files have been imported" but being able to answer four questions reliably: Was the content parsed correctly? Is the text still valid? Can its scope of applicability be filtered? Can a matched fragment be traced back to the complete evidence? If any one of these is missing, it will later show up as "it clearly retrieved the right thing, yet the answer is unreliable."
3. Indexing and retrieval: don't treat semantic similarity as the only answer
The job of the index layer is not to "understand all the knowledge" on the model's behalf but to deliver, with acceptable latency, the text chunks that may contain evidence to the downstream stages. An embedding model maps queries and document chunks into the same vector space, and the vector index then uses an approximate nearest neighbor algorithm to find nearby candidates. What you get here is "possibly semantically related," which is not the same as "able to answer the question."
Choosing an embedding model must be validated on your enterprise corpus
Rankings on general benchmarks are useful only for narrowing down candidates. What an enterprise really needs to test is whether a model can distinguish internal jargon, similar product names, negated expressions, and policy clauses with conditions attached. When selecting a model, check at least the following factors:
- Chinese and industry-specific language: validate abbreviations, aliases, technical vocabulary, and mixed Chinese–English text using real questions, rather than testing only on public question-answering datasets.
- Input length: the length the model can accept should match the chunking strategy. Truncating long chunks outright can leave conclusions, applicability conditions, or exception clauses outside the encoded range.
- Deployment constraints: decide between on-premises and external services based on data residency requirements, inference throughput, latency, and hardware resources.
- Version stability: queries and documents must use compatible embedding models. After a model change, document vectors usually need to be regenerated; you can't upgrade only the query side.
Enterprise retrieval usually needs two retrieval paths
Vector retrieval is good at handling variations in wording. For example, if a user asks "How are travel receipts calculated?", the system may still retrieve content titled "Expense Reimbursement Accounting Rules." But enterprise knowledge also contains large numbers of strings that can't be fuzzy-matched, such as ERROR_CODE_4012, contract numbers, generic drug names, and equipment model numbers. In such queries, a difference of a single character can point to an entirely different object, and keyword retrieval such as BM25 is often more reliable.
| Query characteristics | Preferred retrieval method | Engineering considerations |
|---|---|---|
| Conversational questions, paraphrases, concept descriptions | Vector retrieval | Focus on validating semantic discrimination and the impact of truncating long text |
| IDs, codes, model numbers, proper names | Keyword retrieval | The tokenizer should preserve underscores, hyphens, and complete identifiers |
| Both business semantics and exact terms | Hybrid retrieval | Retrieve candidates separately, then fuse the rankings |
Hybrid retrieval should not be understood as simply adding the two scores together. A safer approach is to generate separate rankings and then use RRF to merge the results based on each candidate's position in the respective lists. If the business has strong constraints on particular fields, rule-based weighting can be added as well, but the weights should be determined through the evaluation set rather than fixed by gut feel.
Metadata filtering should happen before candidates enter the context
In addition to body text, document chunks should carry metadata such as department, product line, publication date, validity status, document version, and access permissions. Filtering by user identity and question scope first, and then running vector or keyword retrieval, keeps repealed policies, versions for other regions, and content the user isn't authorized to see out of the candidate set. Permission filtering in particular can't be handled only at the answer display stage, because once restricted text enters the model's context, the risk of a data leak has already been created.
Filter conditions can't be tightened indefinitely either. The department or time frame may be missing from a user's question, and a wrong inference leads straight to zero results. From an engineering standpoint, distinguish hard conditions from soft ones: access permissions and tenant boundaries are hard filters; product preferences and time tendencies can serve as ranking features, with a step-by-step relaxation strategy designed for cases where nothing comes back.
Top-k should be determined by evaluation results
Too few candidates may miss definitions, conditions, and exceptions scattered across multiple chunks; too many will send the model fragments that are similar in content but differ in scope, increasing erroneous citations and competition within the context. The right range depends on the question type, chunk granularity, index quality, and whether a reranker is configured downstream.
When tuning parameters, watch both "whether the evidence makes it into the candidate set" and "where the correct evidence ultimately ranks." First, gradually increase the retrieval size to measure recall, then set the threshold based on post-reranking accuracy, inference latency, and context usage. If a larger Top-k still doesn't surface the evidence, the problem usually lies not in the number of candidates but in chunking, field filtering, the embedding model, or the keyword index. Using a bigger candidate set to cover up index defects only makes the problem of "finding some relevant content but not reliably answering correctly" harder to pin down.
4. Query processing and reranking: fixing "relevant content was found but not ranked near the top"
Retrieval failures in enterprise question answering don't necessarily mean the knowledge base lacks the material. More often, the relevant fragments have already made it into the candidate set but, because the query is incompletely expressed, the ranking signal is one-dimensional, or the context is poorly assembled, they never reach the model. When troubleshooting this kind of problem, examine the candidate set, the ranking results, and the final context separately; don't judge retrieval quality based only on whether the answer is right or wrong.
Process the query first, but don't overwrite the user's original wording
Users rarely phrase their questions in the formal language of policy documents. Retrieving directly with the original sentence tends to surface content that is lexically similar but has inconsistent constraints.
You can add a query processing layer before retrieval. Common operations include:
- using the conversation history to fill in omitted subjects, departments, or business scenarios;
- mapping internal abbreviations to their official names, while keeping the abbreviations in the matching as well;
- recognizing time conditions such as "this year," "the previous version," and "after the contract takes effect," and converting them into searchable ranges;
- splitting questions that contain multiple conditions into several sub-questions that can each be verified independently;
- generating retrieval expressions closer to the wording of policies, contracts, or technical documents.
Rewritten queries can only serve as additional retrieval entry points; they can't replace the original question. From an engineering standpoint, store the original sentence, the normalized version, the expansion terms, and the split results together, and record which query retrieved each candidate. Otherwise, once the rewriting model misjudges the subject or the time frame, no matter how accurate the downstream ranking is, it will only be picking materials in the wrong direction.
Allocate compute with cascaded ranking
Ranking by vector similarity alone tends to overrate content that is "topically similar"; using keyword retrieval alone can miss synonymous expressions. A sounder approach is to have different retrieval and ranking mechanisms divide the work, rather than asking a single model to make every judgment.
| Processing stage | Main task | Engineering focus |
|---|---|---|
| Candidate expansion | Run vector retrieval and BM25 in parallel, bringing in both semantically related items and exact-term matches | Prioritize avoiding missed retrievals; the candidate count can be relatively generous |
| Fast filtering | Use rules or lightweight models to exclude fragments with the wrong topic, an invalid version, or mismatched permissions | Reduce downstream computation while avoiding premature removal of borderline evidence |
| Fine-grained reranking | Use a Cross-Encoder that reads the question and candidate fragment together to recompute relevance | Focus on whether conditions match, not just on topical similarity |
The value of this cascaded structure is that it reserves expensive judgments for a smaller candidate set. Online parameters should be tuned jointly based on candidate volume, response time limits, and the cost of wrong answers, rather than chasing offline ranking scores alone.
Finishing reranking doesn't mean the context is ready to use
Top-ranked fragments still need to go through context assembly. First, merge chunks with highly duplicated content, so that the same policy passage doesn't appear repeatedly because of copies in different files. Second, for clauses that can only be understood with the surrounding text, add in the heading, definition section, or adjacent paragraphs. Third, arrange materials from the same source in their original chapter order, so that "exceptions" don't come before "conditions of applicability."
You also need to limit the share of the context that any single source can occupy. If the top results are all near-duplicate fragments of the same file, they'll crowd the model's window and keep out another contract appendix or the latest notice that actually determines the conclusion. You can apply quotas by source, document version, and chapter; when conflicts arise, prioritize evidence that has a clear effective date, a matching scope, and directly supports the conclusion.
For multi-condition questions, check evidence coverage rather than just the top score
"If a non-local employee resigns during the probation period, are travel expenses reimbursable, and what approvals are required?" involves the population it applies to, the stage of employment, expense rules, and the approval process. Even if a fragment is highly relevant to "travel expense reimbursement," that doesn't mean the evidence is complete. The system should maintain a coverage status for each sub-question: which have direct support, which have only indirect information, and which still have no material found.
When a single retrieval pass can't cover all the conditions, you can retrieve for each sub-question separately and then merge the evidence according to subject, time, version, and dependencies between rules. Only when evidence coverage meets a predefined requirement should the pipeline move on to answer generation; otherwise it should keep retrieving, ask the user for additional conditions, or state clearly that the available material is insufficient. Ultimately, whether the ranking pipeline is effective is judged not by whether the first result "looks like the answer" but by whether the evidence sent to the model fully supports every constraint the user raised.
5. Answer generation: constraining the model from "answering freely" to "answering from evidence"
Once retrieval is complete, the job of the generation layer is not to have the model "answer using common sense" but to turn the retrieval results into verifiable conclusions. The difference: the former treats the materials as a reference, and the model can still fill in the gaps on its own; the latter treats the materials as the boundary of the evidence, and every key conclusion must be traceable back to specific text.
The prompt should define a clear answering contract: use only the materials currently provided; don't fill in what the materials don't cover from the model's memory; express facts that are directly documented separately from inferences drawn from those facts; and attach the document title, section, page number, or other locating information to every important conclusion. Merely listing a few "reference documents" at the end of the answer isn't enough, because users can't tell which passage actually supports which sentence.
A more reliable approach is to assign each text chunk a stable identifier before it enters the model, and to retain the document version, page number, heading hierarchy, and update time. During generation, require citations to be bound to specific chunks, and after output, have a program check whether each citation ID exists, whether the cited fragment actually contains the relevant facts, and whether key values such as amounts and dates match the source text. Citations should prove the answer, not decorate it.
| State of the material | Generation strategy | Recommended output |
|---|---|---|
| Evidence sufficient and consistent | Organize the answer around the evidence | Conclusion, basis, and item-by-item citations |
| Evidence incomplete | Limit the scope of the answer | What has been confirmed and what information is missing |
| Different versions conflict | Don't choose a version on the user's behalf | List the differences, version dates, and items needing confirmation |
| Beyond the knowledge base's coverage | Stop inferring | Decline explicitly and state what material needs to be added |
The ability to decline is part of generation quality. If the system requires the model to give a definite answer no matter what, the model will tend to fill evidence gaps with linguistic fluency. From an engineering standpoint, design "unable to answer" as a normal state, and distinguish among causes such as no material retrieved, gaps in the material itself, an unclear user question, and insufficient permissions, so that follow-up actions such as adding documents, rewriting the query, or handing off to a human can be taken.
In high-risk scenarios such as policy interpretation, legal consultation, medical information, and financial reconciliation, it's unwise to let the model generate the final response directly from long passages. You can first perform evidence extraction, organizing the applicable subjects, conditions, amounts, dates, ratios, model numbers, and exception clauses into structured fields, and then generate natural language from those fields. Key facts in stable formats should also be double-checked with rules, database records, or calculation logic. The model is responsible for expression; it shouldn't bear the entire burden of fact verification.
How the context is arranged also changes the answer. Correct retrieval results only mean the evidence made it into the candidate set; they don't guarantee the model will use the correct evidence. Before generation, remove duplicate fragments and low-relevance content, place the material that directly answers the question in a more prominent position, and organize the context by document version, topic, or time. If old and new policies appear together, label their effective status explicitly rather than leaving the model to guess.
- Limit irrelevant material in the context so that valid evidence isn't drowned out by noise.
- Prioritize text chunks with intact source boundaries, clear provenance, and valid versions.
- Add version labels to conflicting content and surface the conflict in the output.
- After generation, check citation coverage, consistency of key fields, and unsupported statements.
Answer generation, therefore, is not a single model call appended after the retrieval results but a pipeline that includes evidence arrangement, answer constraints, citation verification, fact checking, and refusal handling. Only by turning "is the answer backed by evidence" into a system rule that can be checked can you truly mitigate the problem of "the material was found, but the answer is still wrong."
6. Evaluation: measure retrieval, generation, and end-to-end business value separately
RAG can't be signed off based only on whether "the final answer looks reasonable." A wrong answer might come from a document parsing failure, from poorly ranked candidate fragments, or from the model ignoring evidence it was already given. Without layered measurement, teams can usually only keep adjusting prompts while the real problem stays upstream.
The first step is to build an evaluation set that supports regression testing. Samples can't be limited to common phrasings; they should also cover low-frequency business questions, questions that can only be answered by combining multiple documents, questions for which the knowledge base has no conclusion, questions constrained by access permissions, and questions whose answers change with version or date. The evaluation set should come as much as possible from real search logs, customer service records, and interviews with business staff, rather than being written entirely by the implementation team.
Each question needs to be labeled with at least the following:
- the acceptable reference answer, as well as the wrong conclusions that must not appear;
- the evidence fragments required to answer, distinguishing core evidence from supplementary material;
- the documents, versions, and effective dates that may be cited;
- the user's identity and access scope, used to verify unauthorized retrieval;
- whether, when knowledge is insufficient, the system should decline outright or may give a limited conclusion.
For answers that are worded differently but mean the same thing, simple string matching isn't appropriate. A more reliable approach is to break the key facts, conditions, values, and scope into check items, and then make a judgment combined with human review.
| Evaluation layer | Key metrics | What to check first when results are off |
|---|---|---|
| Retrieval | Retrieval-quality metrics | Parsing quality, chunk boundaries, index fields, query rewriting, hybrid retrieval, and filter conditions |
| Generation | Factual correctness, evidence faithfulness, citation accuracy, answer completeness, refusal accuracy | Context order, evidence-conflict rules, prompt constraints, model reasoning and instruction-following ability |
| End-to-end | Response latency, cost per call, human handoff rate, resolution rate, user adoption rate | Pipeline complexity, timeout fallback, answer usability, and fit with business processes |
The core of retrieval evaluation is confirming that the necessary evidence reliably makes it into the candidate set. Looking at these ranking metrics alone still isn't enough; you also need to check whether the facts required for the answer are fully covered, and how many fragments that are topically similar but irrelevant to the conclusion have crept into the candidate context.
If the reference evidence wasn't retrieved, don't start by modifying the generation prompt. Trace back along the data pipeline instead: Was the source text parsed correctly? Were table and heading relationships lost? Did chunking separate conditions from conclusions? Did the index include key fields? Did query rewriting change the original meaning? Do keyword, semantic, and metadata filtering complement one another? Check permission filtering and version filtering separately in particular, so that raising recall doesn't bring expired content or unauthorized material into the context.
Generation evaluation should be conducted with the evidence already given. You can put the reference evidence directly into the context to isolate the retrieval variable, and then check whether the answer is consistent with the material, whether the citations really support the corresponding conclusions, whether key conditions were omitted, and whether the system declines to answer when evidence is insufficient. If the material is complete but the output is still wrong, the problem usually lies in context ordering, interference from duplicate information, conflicts between old and new versions, unclear prompt rules, or insufficient model capability.
Citation accuracy needs to be checked item by item; it's not enough to check whether source links appear at the end of the answer. Common failures include citing a passage that is relevant but doesn't prove the conclusion, conclusions and citations drawn from different versions, and the model stitching together multiple documents to produce inferences that don't exist in the source text.
Offline scores must ultimately be validated against business outcomes. Improved metrics may only make answers look more like the reference answers without necessarily reducing the time users spend. Before launch, keep the existing solution as a baseline, run blind human evaluations with the identity of each solution hidden, and then observe real-world effectiveness through small-traffic A/B tests. Beyond accuracy, keep recording tail latency, cost per question-and-answer exchange, the handoff rate to human agents, whether issues are actually resolved, and whether users adopt the answers to continue their work.
The evaluation set should be built into the release process: whenever document parsing, chunking rules, the embedding model, the reranker, prompts, or the generation model change, run the same set of regression tests. New error samples should continue to be added after de-identification and labeling. Only then can you tell which layer an optimization actually improved, and whether it came at the cost of access security, timeliness, cost, or latency.
7. Production governance: keeping knowledge up to date and making sure users see only what they should
After RAG goes live, the main risk is not the model suddenly failing but knowledge versions, access permissions, and index content gradually drifting out of sync. A production system needs to treat document changes as a traceable data pipeline, rather than running a full re-vectorization at regular intervals.
When a new file enters the system, it should go through parsing, chunking, vector generation, metadata writing, and index publication in sequence; when a file is modified or withdrawn, the original text, vector records, keyword index, and caches must be handled in sync. Each chunk should be linked at minimum to its document version, effective period, source location, and processing status. Only after the new version is confirmed to be usable should the old version be marked invalid, preventing two sets of policies from appearing in the candidate results at the same time. Documents whose parsing failed or whose indexing isn't complete can't go live silently; they should enter an alerting and retry queue.
Permission checks must happen at the retrieval stage. Retrieving everything and then letting the model hide sensitive information isn't reliable, because the restricted text has already entered the model's context and may also show up in logs and caches. Retrieval conditions should combine user identity, organizational affiliation, project relationships, document classification, and validity period; after permissions change, the filter data must also be refreshed promptly. Keeping sensitive materials in a knowledge environment controlled by the enterprise helps shrink the data exposure surface, but it can't replace fine-grained authorization.
| Governance object | Required records | Main purpose |
|---|---|---|
| Knowledge changes | Version, effective date, index status, failure reason | Prevent expired content from being used in answers |
| Access control | User attributes, filter rules, permission decisions | Audit unauthorized access and false blocks |
| Question-answering pipeline | Query, candidate fragments, ranking results, final output | Pinpoint which layer an error occurred in |
Online monitoring can't look only at whether API calls succeed. At a minimum, track the share of queries with no results, the share of weakly relevant retrievals, the refusal rate, how citations are accessed, end-to-end latency, and cost per call. User thumbs-downs, expert revisions, and cases handed off to human agents should be categorized by cause—"missing knowledge, chunking error, ranking error, generation drift, permission block"—and turned into evaluation samples for ongoing regression testing.
During implementation, it's best to start with one knowledge domain and one class of stable questions, using basic RAG to establish a measurable baseline. Only when evaluation shows that the bottleneck lies in retrieval, ranking, or multi-hop relationships should you introduce, respectively, hybrid keyword-and-vector retrieval, reranking, query decomposition, or GraphRAG. Stacking too many components at once usually just increases latency and troubleshooting difficulty, and makes it impossible to tell where any accuracy gains came from.
What's the difference between RAG and LLM fine-tuning, and which should an enterprise choose?
RAG solves the problem of querying, updating, and citing external knowledge, and it suits frequently changing content such as policies, manuals, and project materials; fine-tuning is better suited to adjusting output formats, terminology conventions, task steps, or behavioral boundaries. If the problem is that "the model doesn't know about the latest documents," start with RAG; if the model already has the correct evidence but consistently fails to follow a fixed specification, consider fine-tuning. The two can be combined, but fine-tuning shouldn't be responsible for knowledge version management.
Can long-context models replace RAG?
Usually not. Long context increases how much information can be input at once, but it doesn't automatically solve document selection, access isolation, version control, or source tracking. When there's very little material and its scope is fixed, you can load it directly into the context; when materials keep growing, different users are allowed to see different things, or stable citations are required, you still need a retrieval layer to narrow the evidence first. Long context works better as the reading space after retrieval than as a complete knowledge governance solution.
Why is the final answer still inaccurate when the system clearly retrieved the right document?
"Hitting the right document" is not the same as "providing sufficient evidence." The correct fragment may be ranked low, cut off by the context limit, or missing its applicability conditions and exception clauses; the prompt may also allow the model to fill in gaps beyond the evidence. When troubleshooting, check the candidate ranking, the content actually injected, the context length, citation coverage, and the answering constraints separately. The generation stage should require conclusions to be bound to their sources and should decline explicitly when evidence is insufficient, rather than inferring freely.
How many documents do you need to start building an enterprise RAG knowledge base?
There's no universal minimum number of documents. Whether you can start depends on whether the knowledge domain is well defined, whether questions recur, and whether the documents can answer those questions. Structurally stable policy documents can also be used to establish a first baseline. Rather than chasing scale, first prepare a set of real questions, reference answers, and the corresponding evidence; if you can't define this set of samples, importing more documents will only add noise and governance costs.