2026-09-21
What is RAG: a guide to building enterprise knowledge base Q&A
What is RAG? Covering suitable scenarios, data governance, document chunking, hybrid retrieval, reranking, and evaluation, this article systematically explains how to put enterprise knowledge base Q&A into production.
1. First decide whether RAG is the right tool: it solves knowledge supply, not every AI problem
When evaluating RAG, an enterprise should first confirm whether the problem really lies in "the model can't get the right information." The core function of RAG is to select relevant content from an external knowledge source before answering, and then have the model build its conclusions on that content. It is well suited to supplementing internal enterprise knowledge and time-sensitive information, but it can't replace computational systems, business processes, or data governance.
To judge whether a scenario is a good fit for RAG, you can start by checking three conditions. First, the answer depends on private enterprise information, or the required knowledge changes frequently—policy documents, product manuals, project records, and internal case studies, for example. Second, there is documentation against which the question can be verified, rather than relying solely on subjective judgment. Third, the business requires stating where a conclusion came from, and users need to see the original text, version, or publication date. When these three conditions broadly hold, keeping knowledge in an external repository and retrieving it on demand is usually easier to update and easier to trace than baking the knowledge into model parameters.
Conversely, if the source of answers is unclear, the enterprise has never built up the relevant materials, or different documents contradict one another, introducing RAG will only expose knowledge gaps faster. A vector index can help find content, but it can't create facts that don't exist, nor can it automatically determine which of several conflicting documents is authoritative. In that case, first fill in the materials and settle on responsible departments and versioning rules, rather than jumping straight to building a retrieval pipeline.
| Actual need | Better-suited approach | Rationale |
|---|---|---|
| Finding policies and manuals and forming a synthesized answer | RAG | Needs to extract evidence from multiple documents and then generate a question-specific conclusion |
| Returning matching files, pages, or document directories | Enterprise search | The user wants the raw set of results and doesn't necessarily need the model to summarize them |
| Consistently producing a fixed style, format, or behavior over the long term | Prompt templates or fine-tuning | Requires changing how the model expresses itself and its stable patterns, not supplementing factual knowledge |
These approaches aren't mutually exclusive. For example, a customer service assistant can use RAG to look up after-sales policies, read the current status through the order API, and then submit a request through a workflow. The key is to separate "knowledge Q&A" from "deterministic operations": the model can explain the rules, but amount calculations, permission checks, and transaction submissions should be handled by testable programs. Wiring every capability to a vector database tends to produce answers that look complete but can't actually be executed.
Very long context windows don't eliminate the need for retrieval either. Putting large numbers of files into the prompt wholesale seems to remove the retrieval step, but it actually increases input costs and hands the model irrelevant chapters, duplicate versions, and conflicting statements all at once. Key information in long texts may also go underused because of its position and the surrounding noise. For enterprise systems, the safer combination is to narrow the scope of materials first, then use a longer context for cross-paragraph comparison, summarization, and reasoning.
Whether to adopt RAG ultimately comes down to an engineering question: can the answer be found in controlled materials, and is it worth building update, permission, citation, and evaluation mechanisms for those materials? If the question has no verifiable standard answer, design for human judgment or expert review; if a process must be followed strictly, call rule engines, databases, and APIs; if the knowledge source is incomplete, govern the content first. Only when the main bottleneck really is that "the right information isn't getting into the model's context at the right time" is RAG the right remedy.
2. Data governance comes before vectorization: knowledge base quality caps answer quality
An embedding model can only judge whether texts are similar; it can't decide on the enterprise's behalf which document is valid. If the knowledge base holds current policies, historical versions, and unreviewed notes from experience all at once, then the more thorough the retrieval, the more likely conflicting content ends up in the context together. Before launch, resolve questions of knowledge sources, scope of applicability, and lifecycle first, and only then discuss embedding models and vector databases.
Prioritize knowledge sources first
The first step isn't importing files but building an inventory of knowledge sources. Each type of source needs a clearly defined owning department, update method, review status, and conflict-resolution rules. Authority levels shouldn't simply be assigned by file format but determined by business scenario. For example, formal policies usually outrank personal notes; product specifications should follow the released version; FAQs can explain standard documents but shouldn't override their constraints; tickets are good for supplementing case examples but shouldn't automatically become general conclusions.
| Knowledge type | Main use | Conflict-handling principle |
|---|---|---|
| Policies and standards | Define rules, permissions, and boundaries | Prefer the currently effective, approved version |
| Product and operations documentation | Answer questions about features, configuration, and processes | Match by product version, region, and release date |
| FAQs and knowledge articles | Provide simplified explanations and common solutions | Must not rewrite formal rules; downweight or block on conflict |
| Tickets and personal records | Supplement incident cases and field experience | Treated as reference material by default; promoted only after review |
Metadata isn't decoration; it's a retrieval boundary
Field values should use controlled enumerations to avoid the same department or product being written in multiple ways. This information is used for pre-retrieval filtering on the one hand, and supports citation display, access authorization, and expiry management on the other.
Otherwise, a regulation with highly similar wording but applicable to a different market may well rank ahead of the correct answer.
Preserve document structure and evidence locations during ingestion
Enterprise documents can't simply be extracted into continuous plain text. A reliable ingestion pipeline usually needs to handle the following steps:
- Identify the body area and strip out repeated headers, footers, watermarks, and table-of-contents noise;
- Preserve section headings and their parent-child hierarchy so that each chunk can still explain the context it sits in;
- Parse the row and column relationships of tables, and link the attachments, figures, and footnotes referenced in the body;
- Detect duplicate content across files or pages so the same conclusion doesn't crowd the retrieval results;
- Record file identifiers, page numbers, paragraph paths, and character ranges so you can go back to the original to verify.
Parsing quality should be verified through sampling-based acceptance, not just by whether the job succeeded. Focus on checking whether tables are misaligned, headings are lost, or attachment links are broken, and whether negative conditions and exception clauses have been split apart. Keeping only the semantic text while losing the original location makes citations appear to exist when in fact they can't be audited.
Keep index state in sync with changes in source systems
RAG can update external knowledge without adjusting model parameters, but "updatable" doesn't mean "naturally consistent." The ingestion process must cover additions, revisions, expirations, withdrawals, and deletions, and save the source identifier, processing time, and index state for every change.
On the engineering side, you should also set up monitoring for sync delays, parsing failures, coexisting versions, and orphaned chunks, and periodically cross-check the index against the source systems. There's a direct standard for judging whether data governance is up to par: every answer chunk can state where it came from, whom it applies to, and when it is valid, and it stops being retrieved within the agreed time after the source document is withdrawn.
3. Chunking and index design: don't handle every document with one fixed parameter
Chunking isn't about cutting long text down to a length the model can accept; it's about constructing knowledge units for retrieval that can be judged independently and cited accurately. The real question to optimize is: when a user asks a specific question, can the system retrieve a chunk that contains the complete conditions, conclusion, and scope of applicability? If you just truncate at a fixed character count, the exceptions in a clause, the prerequisites of an operating step, and the meaning of fields in a table may all get split apart.
When chunks are too large, retrieval results may hit the topic but send large amounts of irrelevant text into the model along with it, diluting the relevant evidence. When chunks are too small, the retrieved content looks precise but lacks conditions, exceptions, or conclusions, leaving the model to fill in the gaps on its own. Overlap between adjacent chunks can reduce the problem of sentences landing right on a boundary, but it can't fix a wrong document structure. For example, if the main text of a policy provision and its proviso are split into two chunks, even with overlap there's no guarantee both will make it into the candidate set.
In engineering practice, teams often set up several candidate chunk lengths on the order of a few hundred tokens, configure some overlap between adjacent chunks, and then compare results using a real question set. These parameters are only an entry point for experimentation, not a universally optimal configuration across documents. A more robust implementation first identifies heading levels, lists, clauses, Q&A pairs, and tables, then performs a second split on semantic units that are too long; chunks that are too short can be combined with a parent summary or heading path for retrieval.
An index also shouldn't equal "text plus vectors." Each knowledge chunk should store at least the following information:
- The body text for the generation model to read, plus normalized text for retrieval;
- The document title, section hierarchy, and parent-child chunk relationships;
- Stable document identifiers, chunk identifiers, and the original page number or paragraph position;
- Publication time, update time, validity status, and business version;
- Permission labels such as department, role, and classification level;
- Keywords, entities, or other fields usable for filtering and keyword search.
Vectors are good at finding semantically similar content, but they aren't necessarily reliable for exact identifiers, product models, abbreviations, dates, and proper nouns. So an inverted index over the body text, structured filter conditions, and vector fields should coexist. Heading paths can also fill in context during generation—for example, chunks titled "Approval conditions" mean different things in a procurement policy and an expense policy.
Finally, the index must support rebuilds and rollback. Each release should record the source document batch, parser version, chunking rules, text-cleaning logic, Embedding model, and index configuration, and preserve the mapping from chunks back to the original text. Only then, if recall drops after a model upgrade, can the team determine whether the problem stems from document parsing, changed chunk boundaries, changed vector representations, or index filter configuration. When a knowledge base without version records fluctuates in quality, it usually can only be rebuilt from scratch, making verifiable fixes difficult.
4. From user question to effective retrieval: hybrid search for "can't find it" and "finds the wrong thing"
The online pipeline of enterprise RAG shouldn't start directly with a vector query. User input often isn't a proper search query: it may omit the product name, describe an error in colloquial terms, carry over references from the previous turn of the conversation, or mix in implicit conditions such as version, region, or job role. If the raw question is sent to the index unprocessed, then even increasing the number of retrieved results later will only yield more text that looks similar on the surface but doesn't actually apply.
The first step should be to turn conversational input into a question that can be understood on its own. The query processing module needs to extract the task the user actually wants to accomplish and fill in the object scope, applicable version, time conditions, and user identity. For example, "how do I handle this error" should be reconstructed, using the conversation history, into a specific product, version, and error code. Abbreviations, internal aliases, and colloquial names can be mapped to standard terms in the knowledge base, but the original words should also be kept to avoid losing precise clues during rewriting.
Complex questions shouldn't be forced into a single search query. If a question asks you to look up rules, find data, and compare differences all at once, it can be split into several subqueries retrieved separately, with the evidence merged at the generation stage. The value of splitting isn't making the question shorter but keeping multiple retrieval goals from interfering with each other. In implementation, preserve the mapping between subqueries and the original question; otherwise the model may answer only part of it.
| Retrieval method | What it hits best | Typical risk |
|---|---|---|
| Vector search | Synonymous expressions, natural-language descriptions, passages with different wording but similar meaning | May rank content that is semantically similar but for a different version or object near the top |
| BM25 keyword search | Device models, error codes, contract terms, statute numbers, and enterprise-specific terminology | Easily misses results when users phrase things differently, and struggles to understand contextual meaning |
| Hybrid search | Covers both semantic cues and exact literal cues | If fusion weights are fixed, it may still be ill-suited to different types of questions |
So an enterprise knowledge base should usually run vector search and BM25 in parallel and then merge candidates with rank fusion, rather than choosing only one. When a query contains an error code, model number, or clause number, you can increase the weight of keyword results; for natural-language questions such as fault descriptions or process inquiries, you can lean toward semantic retrieval. You should also deduplicate candidates so that adjacent chunks from the same document don't fill up the context.
The user's organization, access level, business region, product version, and document validity period should all become index metadata. At query time, first restrict the available scope, then compute relevance within that scope. Otherwise, the system may retrieve files whose content is correct but which the user isn't authorized to see, or cite regulations that have been repealed, aren't yet in effect, or apply only to other regions. Merely asking the model in the prompt "not to leak" information is no substitute for access control.
- Checking "can't find it": has the target evidence been indexed, was the query rewritten incorrectly, and did at least one of the keyword and semantic channels get a hit?
- Checking "finds the wrong thing": do the top-ranked results satisfy the product, version, region, timeliness, and permission conditions, rather than just text similarity?
- Recording the retrieval trace: save the original question, rewrite results, subqueries, filter conditions, and the rankings from each channel so failure cases can be reproduced.
For questions that require following multiple entity relationships to find evidence, or that ask for overall themes to be summarized from large numbers of documents, GraphRAG is worth evaluating. It organizes knowledge by entities and relationships and is better suited to cross-document relationship chains and global analysis, but the costs include entity disambiguation, relationship updates, graph query maintenance, and a longer response path. In engineering terms, first use a question set to verify where ordinary RAG falls short: if failures mainly come from query rewriting, filter conditions, or hybrid search configuration, building a graph structure won't solve the root cause. Only when the core questions truly depend on multi-hop relationships or global structure may the added complexity be worth it.
5. Reranking and answer generation: avoiding "retrieved the right material, still answered the wrong question"
A retrieval hit doesn't mean the evidence is usable. First-stage retrieval is more like "widening the search": better to bring back a little noise than to miss key material too early. What should really decide which content enters the model's context is the subsequent reranking. A Reranker reads the user's question and each candidate chunk together and re-judges whether they match semantically, rather than continuing to rely on vector distance.
When the enterprise corpus is large, it's unwise to have a computationally expensive fine-ranking model scan everything. A more controllable approach is cascaded processing: first use keyword search, vector search, and similar methods to quickly obtain a broad candidate set, then use a lightweight model to eliminate obviously irrelevant items, and finally perform fine-grained ranking with a Cross-Encoder-type model. Each layer should record candidate counts, latency, and reasons for elimination, making it easier to tell at which step quality was lost.
Reranking can't just look at "similar topic." For example, if a user asks about the conditions under which a certain policy takes effect, a passage introducing the policy's background may be highly semantically relevant yet unable to answer what the conditions are. Scoring should additionally take the following factors into account:
- Whether the content can directly support the conclusion the question calls for, rather than merely sharing the same topic;
- Whether the material is within its validity period, and whether version, region, and applicable audience match;
- Whether the source tier is reliable—formal policy documents, for example, should take precedence over meeting minutes or personal notes;
- Whether the chunk contains the complete conditions, exception clauses, and scope limits, to avoid picking up only half a conclusion.
Before entering the generation stage, the context also needs to be orchestrated. Mechanically concatenating chunks by relevance score often causes the same passage to appear repeatedly, separates clauses from their headings, or places explanatory text before the main text of a rule. In engineering terms, first remove duplicate chunks, merge consecutive chunks that belong to the same section, and then add back headings, document versions, publication dates, and whatever preceding text is needed to understand the current passage. If the token budget is tight, prioritize material that answers directly, has a clear source, and is current, rather than compressing all candidates evenly.
The generation prompt should be treated as a testable answer protocol, not a single line saying "please answer based on the materials." For conclusions that affect approvals, compliance, or business execution, attach the original text excerpt and its document location so users can go back to the source to double-check.
Citations themselves also need validation. The citation numbers the model generates must correspond to materials that actually entered the context, and the cited excerpt must genuinely support the adjacent conclusion, not merely establish relevant background.
When a wrong answer appears, first locate which layer failed, then decide what to change. Attributing every problem to "the LLM isn't strong enough" usually leads to ineffective parameter tuning.
| Failure type | How to identify it | Main remedial actions |
|---|---|---|
| Knowledge gap | Manual inspection of the knowledge base confirms that no material exists to support the answer | Add authoritative documents, fix version and permission scopes, and establish a refusal policy |
| Retrieval failure | The answer exists in the knowledge base, but the candidate results don't include the corresponding chunk | Adjust query rewriting, chunking, index fields, hybrid search, and candidate scope |
| Evidence-use failure | The correct material is already in the context, but the answer still contradicts it or omits limiting conditions | Improve context orchestration and prompt constraints; add citation validation and faithfulness evaluation |
So acceptance at this stage shouldn't just check whether the final answer "looks right." You also need to check separately whether the correct evidence entered the candidate set, whether fine ranking lifted it into a visible position, whether the context preserved its full meaning, and whether each conclusion can be traced item by item back to the evidence. Only by observing this pipeline piece by piece do off-target answers turn from sporadic user-experience problems into engineering problems that can be located and fixed.
6. Answer evaluation and the operations loop: measure retrieval, generation, and business outcomes separately
RAG evaluation can't start from questions the model makes up for itself. Besides the reference answer, each sample should also be labeled with the evidence needed to answer it, the product or time range the answer applies to, access permissions, and whether the question should be refused. Otherwise, the system may give answers whose content is correct but whose version is wrong, that exceed the user's authorization, or that are missing prerequisites—without the composite score revealing any problem.
Coverage matters more than the number of samples. At a minimum, include abbreviations and aliases, typos, omitted context, multi-turn follow-ups, conflicts between old and new policies, permissions for different roles, and questions that can only be answered by combining multiple documents. You should also build a dedicated unanswerable set—for example, content the knowledge base hasn't yet included, questions with insufficient conditions, contradictory evidence, or content the user isn't authorized to see. Refusal isn't an exception branch; it's a basic capability of enterprise Q&A.
| Evaluation layer | Question to answer | Metrics to record separately |
|---|---|---|
| Retrieval | Was the necessary evidence found, and did it make it into the context the model actually reads? | Evidence recall in the candidate set, hit rate in the final context, share of irrelevant chunks, correctness of permission filtering, version selection results |
| Generation | Did the model answer fully based on the evidence, rather than filling in plausible-sounding content? | Factual correctness of answers, faithfulness to evidence, citation location accuracy, coverage of key points, behavior when refusal is required |
| Business | Has the system reduced the cost for users of finding information and of manual inquiries? | Hand-off-to-human rate, no-answer rate, problem resolution, user feedback, repeated follow-up questions |
| Engineering | Do quality gains come with unacceptable resource overhead? | End-to-end latency, model token usage, reranking overhead, failure rate, and stability at peak |
If the correct evidence doesn't make it into the candidate set, you should usually check query rewriting, indexing, or the retrieval strategy; if the evidence was retrieved but didn't make it into the final context, the problem more likely lies in reranking or context assembly; if the context is complete but the answer is wrong, keep checking prompt constraints, the citation mechanism, and model generation. Only with layered records can failures be attributed, rather than lumped together as "the model isn't performing well."
Citations should also be verified separately. An answer that comes with sources doesn't mean the sources support its conclusions. During evaluation, check whether the cited excerpt really contains the corresponding fact, whether it points to the currently valid version, and whether a single citation is being wrongly used to support multiple conclusions. The value of being able to pinpoint specific knowledge chunks isn't only that users can verify them easily; it also lets the team trace back to source materials that are outdated, conflicting, or ambiguously worded.
Offline evaluation is for comparing approaches—for example, observing how different types of questions change after adjusting chunk boundaries, swapping retrieval methods, or modifying ranking rules. Online data is for judging whether the system has produced business benefits. Neither can replace the other: improved offline hit rates don't mean users are solving problems faster, and a drop in hand-offs to human agents may just mean users have given up asking. So analysis should combine session completion, follow-up questions, and negative feedback.
The core of the operations loop is building a failure queue with clear attribution. Missing or conflicting knowledge flows back to content governance; misunderstood phrasings go to query rewriting; truncated meaning goes to chunking optimization; target material that never appears goes to retrieval tuning; relevant material ranked too low goes to reranking; and answers that go off track despite having evidence go to context organization and prompt strategy optimization. Each case should retain the question, permissions, knowledge version, retrieval results, final context, and model output, rather than saving just a snippet of the wrong answer.
Before any change goes live, run a fixed regression set and review differences by question type. A local parameter tweak may fix one kind of phrasing while degrading old-version recognition, permission isolation, or multi-document answers. Mature RAG operations isn't about pushing up a single score; it's about making sure every failure has an owner, every change can be verified, and every release tells you what improved and what it might have affected.
7. FAQ: common questions when building RAG in the enterprise
With very long context windows, do we still need RAG?
Yes, but the two aren't substitutes for each other. A longer context window addresses how much material the model can take in at once; RAG addresses which material from enterprise knowledge is worth sending to the model. Stuffing large numbers of documents directly into the prompt can increase the processing burden and hurt permission management and evidence use. Information in the middle of long texts may also be harder for the model to use reliably—what's often called "lost in the middle."
The safer approach is to use RAG first to narrow the scope of evidence, then let the long context handle cross-chapter reading, multi-document comparison, and complex synthesis. Using long context directly may be simpler only when the document set is small, the content is relatively fixed, and a single task truly requires reading everything. For example, contract materials with a clearly defined scope of analysis can be input in full; answering clause questions across a large contract repository should start with retrieval.
How should you choose between RAG, fine-tuning, and enterprise search?
First distinguish whether what you want to change is knowledge, behavior, or the way information is accessed:
| Approach | Problems it suits | Main limitations |
|---|---|---|
| RAG | Knowledge updates often, answers must cite internal materials, and access control must be enforced | Depends on document governance, retrieval quality, and generation constraints |
| Fine-tuning | Stable output formats, terminology style, classification rules, or specific task behaviors are needed | Not suited to serving as a repository of frequently changing facts; updates require preparing data and retraining |
| Enterprise search | Users want to browse the original text, filter results, and make judgments themselves | Usually delivers only a list of results and doesn't synthesize multiple sources into an answer |
Real projects often combine them: search handles discovery and navigation, RAG generates explanations based on evidence, and fine-tuning constrains the model's task behavior. Don't jump straight to fine-tuning just because "answers are inaccurate"; if the root cause is missing material or retrieval failure, training can't bring back the correct evidence.
Vector search found similar content—so why is the answer still wrong?
"Semantically similar" doesn't mean "usable for answering." A retrieved chunk may share the topic but not match the question on product version, applicable region, customer tier, or effective date; or retrieval may return only the conclusion while missing definitions, exception conditions, and context. Another class of problems arises at the generation stage: the evidence is there, but it's ranked too low, gets drowned out by irrelevant chunks, or the model stitches together multiple sources incorrectly.
When troubleshooting, inspect the pipeline layer by layer rather than reading only the final answer:
- First confirm whether the correct document made it into the candidate set; if not, check chunking, query rewriting, keyword retrieval, and metadata filtering.
- Then confirm whether the correct chunk ranks in a position the model can actually see; if ranking is poor, introduce reranking and reduce the share of duplicate chunks.
- Finally, check whether the answer is strictly constrained by the evidence: require sources, distinguish facts from inferences, and refuse explicitly when evidence is insufficient.
What metrics must enterprise RAG be evaluated on, at minimum, before launch?
Acceptance can't rest on "the answers look pretty good." Build a test set from real business questions and break the metrics into three layers:
- Retrieval layer: whether the correct evidence makes it into candidate results, whether it ranks near the top, whether filter conditions such as permissions and time are accurate, and whether retrieved content contains a lot of duplication.
- Generation layer: whether answers are supported by evidence, whether citations map to the original text, whether the question is fully covered, whether the system can refuse when evidence is insufficient, and whether unsupported additions appear.
- Business layer: task completion rate, hand-offs to human agents, how often users make corrections, response latency, and cost per call.
After launch, keep collecting low ratings, follow-up questions, and records of manual rewrites to determine whether problems originate in the knowledge source, retrieval, ranking, or prompting and generation. Only when you can pinpoint which step an error occurs in does RAG become operable, rather than a one-off demo system.