2026-06-27
Building a knowledge base Q&A system: architecture selection and deployment paths
Building a knowledge base Q&A system that is genuinely "usable" takes far more than getting a demo to run. Starting from five acceptance criteria, this article covers retrieval quality, multimodal routing, a knowledge-first strategy, and observability, then compares where each of three deployment paths fits and what it trades off: a custom-built RAG stack, an open-source product, and low-code orchestration. The goal is to help teams avoid detours when making architecture decisions.
Why a knowledge base Q&A system that "runs" is not the same as one that is "usable"
Most knowledge base Q&A demos go smoothly: upload a few documents, ask a question, the model returns a plausible answer, and the room nods its approval. The problem is that an engineering chasm separates this scenario from production. A demo validates that "this pipeline can produce an answer," while production has to answer a different question: "can this pipeline consistently produce answers that are trustworthy, controllable, and accountable?" The two are not in the same league of difficulty.
What really decides whether a project succeeds is not whether the principles of RAG have been explained clearly. The basic logic of retrieval-augmented generation (turn the question into a vector, find similar passages in the document store, insert those passages into the prompt, and have the model answer) is not complicated; online tutorials can get an engineer to a working prototype in a day. But between a working principle and a usable system lies a long list of engineering problems that are invisible in demos and erupt all at once after launch. What does the model do when retrieval occasionally comes back empty? How do you ensure answers are trustworthy when the retrieved passages vary widely in quality? What happens when a user asks about data in a table but the system only retrieved body paragraphs? None of these are questions of principle. They are engineering problems.
For decision-makers, the question to answer has never been "how is RAG implemented?" That is for the technical team to work out. What decision-makers actually need to judge is something else: will this system fail after launch, in which scenarios will it fail, and who bears the cost when it does? A Q&A system that fabricates answers is more dangerous than having no Q&A system at all. It leads business staff to make decisions based on wrong information, and because the answers look professional, the errors often go unnoticed until they have caused real consequences. Usability is not a vague impression; it is a set of engineering conditions that can be checked one by one.
That is why this article does not start from the principles. There is already plenty of content on them, and writing it up again would neither add anything new nor help anyone make a decision. We take a different angle instead: we break "is this system usable?" down into five acceptance criteria. Each criterion maps to specific configuration items, settings that can be inspected on site, and clear failure signals that appear when it is not met. The acceptance logic works in reverse: rather than arguing how the system should be designed, it lists the symptoms that will surface if the design is flawed. With this checklist in hand, decision-makers can bypass the technical details and check a vendor's proposal or an in-house team's work directly against it.
The five criteria cover the stages where things most often go wrong. First, whether the model's answers are strictly confined to the retrieved content, or whether it improvises freely when retrieval comes back empty. Second, whether the number of retrieved results and the quality threshold for retrieval are adjustable parameters rather than a hard-coded black box. Third, when the knowledge base contains non-text content such as images, tables, and scanned documents, whether the multimodal processing pipeline is triggered correctly or silently skipped. Fourth, whether the system actually queries the knowledge base first, or frequently bypasses it and answers from the model's general knowledge, leaving the materials you carefully maintain unused. Fifth, whether the whole system can be observed, maintained, and integrated into existing workflows, rather than being an island nobody dares to touch.
These five are not a theoretical framework; they are checkpoints worked backward from launch failure after launch failure. Behind each one is a class of failure that has actually happened. Use them as an acceptance checklist: a system that passes every item is very likely to hold up in production; if any item gets stuck, fix it before launch rather than waiting for users to discover it for you. The sections below go through each criterion in turn, explaining what to check, how to check it, and what you will see when it falls short.
Acceptance criterion 1: retrieval results are non-empty, and the model answers only from retrieved content
This is the most basic criterion and also the one most often skipped. Many teams test only whether "Q&A produces a result" before launch, without distinguishing two things: a model answering fluently and a model answering with a grounded basis are entirely different matters. A system that can spin a story looks more pleasing than one that honestly admits "I don't know," but in production the former is a ticking time bomb.
Start with empty retrieval. If you find that some questions clearly match content in the knowledge base but the system returns a blank or an irrelevant answer, your first move should not be tuning the similarity threshold. Go back to the knowledge base management interface and confirm the indexing status of those documents. A successful upload does not mean a document is retrievable; it still has to go through chunking and vectorization before it enters the retrieval pool. This process is asynchronous and can take several minutes for large files or bulk imports. I have seen far too many "retrieval returns nothing" tickets that turned out to be documents stuck in a processing state that had been treated as live. So the hard requirement before launch is to verify, document by document, that processing has completed for everything that should be retrievable, rather than glancing at the upload list and signing off.
The second issue is subtler and concerns what the model is actually answering. The retrieval step may have pulled out the relevant passages, but those passages have not necessarily made it into the model's view. In most orchestration frameworks, the retrieval node produces a result variable that holds the retrieved text. If you do not attach that variable to the context in the LLM node, the model receives only the user's original question, so it answers freely from its own pretraining knowledge. On the surface everything looks normal, and the answer even seems reasonable, but the model never looked at your knowledge base. The correct approach is to inject the retrieval output explicitly into the context and reference it with a context placeholder in the system prompt, telling the model clearly that "you may answer only based on this material." Get this step wrong and the entire RAG pipeline degrades into an ordinary chatbot, with the knowledge base serving no purpose.
These two problems share an inconvenient trait: they don't throw errors. The system keeps running, answers keep being generated, and the monitoring dashboard stays green. The problems surface only when you check real business questions against reference answers, and by then it is often after users have already complained.
So to pass this criterion, checking that normal questions work is not enough; you have to test in reverse. The most effective step is to deliberately ask a question whose answer does not exist in the knowledge base, for example by inventing a nonexistent product model or asking about a policy detail outside the scope of the materials. Then see how the model responds:
- A passing answer states clearly that "no relevant content was found in the available materials," or guides the user to rephrase the question or provide more information.
- A failing answer confidently delivers a passage that sounds professional but is entirely fabricated. This is classic hallucination, and hard evidence that the model is improvising beyond the knowledge base.
Turn this test into a fixed set of "trap questions" and run it every time you change the prompt, swap models, or adjust retrieval parameters. It is more reliable than any subjective impression. A system should say "I don't know" a few more times rather than talk nonsense to users with a straight face. In enterprise knowledge scenarios, the cost of a wrong answer is far higher than the cost of no answer.
Passing this criterion shows that your retrieval pipeline works and that the model is genuinely constrained to the boundaries of your knowledge. But "can answer from the knowledge base" is only a passing grade; whether answers are accurate and complete depends on the quality of retrieval itself. The next criterion deals with that.
Acceptance criterion 2: retrieval quality and the number of retrieved results are controllable and tunable
The first criterion addresses "whether the answer has a basis"; this one addresses "whether that basis is found accurately and supplied in sufficient quantity." The dividing line for when a Q&A system becomes truly usable often lies not in how powerful the model is, but in whether the retrieval layer exposes enough tuning knobs. If retrieval behavior is a black box, with a fixed number of results, no choice of mode, and invisible chunks, then when the system starts giving irrelevant answers after launch, you won't even have a place to start tuning.
Start with the retrieval mode. Pure vector retrieval is good at capturing semantic similarity but often misses in cases that require exact matches, such as proper nouns, product model numbers, and error codes; pure keyword retrieval is the opposite, matching words but not meaning. In practice, the safer approach is to have the retrieval node run both semantic and keyword search and then fuse the results, commonly known as hybrid retrieval, with reranking layered on top to improve the relevance of the top results. The output of the retrieval node should not be a single block of concatenated plain text but a structured array of objects, each carrying a document chunk along with metadata such as its source and score. This is critical: structured output means the downstream LLM step receives clear context boundaries, and you can filter, truncate, or rerank each retrieved result within the pipeline, instead of facing a string that cannot be taken apart.
Next, the number of retrieved results. How many documents each retrieval round returns is fundamentally a trade-off between recall and context length. Too few, and the relevant information may never make it into the context, so the model naturally gives incomplete answers; too many, and irrelevant chunks dilute the useful signal while eating into the token budget and driving up latency and cost. A reasonable engineering default is 3 results per round, with the upper limit left open for users to adjust by scenario. Short FAQ-style Q&A needs fewer results, while analytical questions that synthesize multiple passages may need a correspondingly higher number. The point is not some magic number; the point is that this value must be configurable, so it can be tuned back based on actual Q&A performance rather than hard-coded.
What is truly easy to underestimate is the chunking strategy, which sets the ceiling on retrieval quality at the ingestion stage. A common default configuration is 1000-character chunks with a 200-character overlap, which is usually adequate for loosely structured long text. But defaults are a fallback for general cases, not a fit tailored to your documents. The problem is that fixed-length chunking can easily cut right through the middle of a semantic unit: a table split in two, a code example separated from its beginning or end, a complete set of instructions truncated at step three. Once these fragmented chunks enter the vector database, each one is semantically incomplete on its own; they are hard to retrieve, and even when retrieved they cannot support a complete answer from the model.
So always spot-check the chunking results after ingestion. Open the chunks that were actually generated and take a look. A medium-length Markdown document will typically be split into a dozen or so chunks; confirm one by one that they fall on reasonable semantic boundaries. If you find that strongly structured content such as tables, lists, and code blocks is frequently cut apart, go back and adjust the chunk granularity, or introduce a strategy that chunks by heading hierarchy or paragraph structure so that chunking respects the document's own organization. This work looks tedious, but it offers the highest return of anything in retrieval quality. It sits furthest upstream, and once chunking is broken, no amount of tuning of modes or result counts downstream can make up for it.
To turn this criterion into acceptance checks: the retrieval mode can be switched and hybrid retrieval is enabled by default; the number of retrieved results is configurable with a reasonable default; and chunking results have been manually spot-checked to confirm that no semantic units are broken apart. Only when all three are met is the retrieval layer handed over in a state that can be continuously tuned, rather than as a black box nobody can adjust.
Acceptance criterion 3: the multimodal pipeline is routed and triggered correctly
The first two criteria govern the reliability of a single text pipeline, but in real scenarios users don't send only text. An ops colleague will screenshot an error log, after-sales staff will photograph an equipment nameplate, and finance will toss over a scanned invoice. Whether the system can recognize that "this input includes an image" and send the request down the right processing channel is the core of the third criterion. The test is straightforward: in the same chat box, with the user changing nothing, plain-text messages go through text retrieval, messages with an image attached go through visual understanding, and both produce reasonable answers. If the system can't do this, users are forced to remember "use this entry point for text questions and that one for images," and the experience falls back into a fragmented state.
The key to implementation is the routing logic. When a request comes in, the first step is to determine whether this turn includes an attached file. Requests with files are routed to a vision model that can read images; requests without files go through the text-only retrieval-augmented channel, where the language model answers with the knowledge base context. The check itself is not complicated. The hard part is that it must sit at the very front of the pipeline and have a clear destination for both kinds of input, so that neither "image requests landing in the text channel" nor the reverse can happen. If the routing check is written loosely, for example handling only the branch with a file and missing the fallback for no file, plain-text questions can get stuck at a node waiting for an image that never arrives, which shows up as inexplicable timeouts or empty responses.
The Vision toggle on the vision node: the setting most often missed
One extremely common configuration oversight deserves to be called out on its own. Vision model nodes usually come with a toggle that controls whether image input is processed, and it is often off by default. If the routing itself is set up correctly and the image really does reach the vision node, but the toggle was never switched on, the result is not an error exit. Instead, the model treats the image as unreadable content and replies with something like "unable to recognize the file." The trouble is that this feedback looks too much like a "normal failure": whoever configured it assumes the model isn't capable enough or the image format is wrong, then spends ages swapping models and compressing images, when the root cause is simply an unchecked option.
The reason to stress this repeatedly is that the cost of troubleshooting this fault is wildly out of proportion to its actual complexity. Work backward from the symptoms: if uploading an image consistently returns "unable to recognize" while text questions work fine, the problem is most likely not the routing but that the vision node isn't reading images at all. Checking the toggle first is much faster than suspecting the model and the image first. Add this to your team's troubleshooting checklist and you will save a great deal of pointless trial and error.
How to accept it: run each pipeline once
The simpler the acceptance steps, the better, because they need to be run repeatedly: every time you change the routing, swap a model, or adjust a node's configuration, run them again. Specifically, do two things:
- Ask one text-only question. Pick a question whose answer definitely exists in the knowledge base, and confirm that the answer comes from the knowledge base and that no image-related processing was triggered. This verifies that the no-file branch is hit correctly.
- Ask one question with an image. Upload a clear, business-relevant image along with a text request, and confirm that the vision node actually read the image and answered based on its content, rather than replying "unable to recognize." This verifies that the file branch is hit and that the Vision toggle is really in effect.
Only when both pass is the pipeline truly working end to end. There is an easily overlooked point here: many people test only the image path, because vision is the new feature and the one that worries them most, and assume plain text has been fine all along. But routing is an either-or decision, and changes to the image branch can easily break the no-image branch, so both sides must actually be run, not inferred.
There is also a more advanced acceptance angle: try a question "with an image that is irrelevant" and see whether the system falls back to plain-text retrieval or insists on finding an answer in the image and ends up answering off-target. This step is not mandatory, but it can reveal whether the routing logic is too crude, checking only "is there a file?" without considering whether the file actually needs to be understood. For scenarios where text and images are frequently mixed, thinking through how to handle such edge cases in advance is far better than being caught off guard by real user input after launch.
Get these three things right and the multimodal part is basically stable. Unlike retrieval quality, it doesn't require continuous parameter tuning; it is mostly a matter of configuring routing and toggles correctly once and then guarding against regressions with a fixed set of checks. It is the kind of work that "pays off for a long time once it's set up right."
Acceptance criterion 4: knowledge retrieval comes first, ensuring the knowledge base is actually used
One type of failure is very well hidden: the system runs and its answers look reasonable, but when you go back through the logs, you find that the share of requests that actually hit the knowledge base is absurdly low. You built the knowledge base and spent the money, yet most answers come from the model's own parametric memory and have little to do with your documents. This kind of "bypass" throws no errors, so if you don't specifically check for it during acceptance, it is easy to miss.
The problem lies in the routing order. A common flawed design first determines the type of user input (whether a file was uploaded, whether it's an image, whether web search is needed) and then decides which pipeline to take. Once this branching logic handles "plain-text questions" and "questions with attachments" separately, knowledge retrieval is often attached to only one of the paths, leaving the other one running unprotected. A user casually drops in a screenshot, the system feeds it straight to the vision model, and the knowledge base is never touched.
The more robust approach is to move knowledge retrieval ahead of all branches as an unconditional first step. Whether the input is text, a file, or an image, first use the user's intent to pull relevant chunks from the knowledge base, attach the results as reference context, and only then proceed to the type check. That way, no matter how the downstream pipeline branches, the model always has material from your knowledge base in hand. Empty retrieval is a separate matter (covered by acceptance criterion 1); what needs to be guaranteed here is that "the retrieval step always happens," not that it is triggered on a whim.
The vision pipeline is where this principle is most often overlooked. Many teams assume images should be handed to a multimodal model to look at in isolation, and answered directly afterward. But in reality, users who upload images usually bring context with them: a screenshot of an equipment error, accompanied by the question "how does our repair manual say to handle this fault?" If the vision model only describes what is in the image and then answers from nothing, what it gives is general common knowledge, not the standard procedure in your manual. The correct approach is to feed image and text together: the vision model parses the image content while knowledge retrieval results are injected at the same time, so the model can combine "understanding this image" with "checking it against our materials." During acceptance, you can build a dedicated set of "image + knowledge-base-related question" test cases and check whether the answers cite information from the knowledge base, rather than stopping at an objective description of the image.
The third aspect is extensibility. Today your knowledge base may contain only text and images; tomorrow you'll need to add PDF parsing, audio transcription, and table extraction. If the architecture didn't leave room for these branches from the start, every new input type will require changes to the core routing, and every change will break something. A more pragmatic design makes input processing a set of pluggable branches, so that when a new format arrives you simply register one more preprocessing pipeline, with no changes to the shared step of retrieving knowledge first or to the downstream model calls.
When accepting a new branch, the focus is not on whether the new feature itself runs, but on whether plugging it in has broken the old pipelines. Keep a set of regression cases that cover the existing input types, and rerun the entire set after each extension: do the existing plain-text Q&A and image-and-text Q&A still hit the knowledge base as expected, and has the routing wrongly diverted requests to the new branch? A common failure scenario is a newly added PDF branch capturing traffic that should have gone through ordinary text handling, or audio transcription output bypassing the retrieval-first step. None of these are feature bugs; they are integration bugs that unit tests can't cover and that only surface through end-to-end regression runs.
To condense this criterion into one executable check: randomly sample a batch of real production requests, calculate the share that actually triggered knowledge retrieval, and then see how many of the image and file requests used knowledge base content in their answers. If the share is clearly low, your knowledge base is being quietly sidelined by the routing, which is more worrying than wrong answers. You can see wrong answers; you can't see being bypassed.
Acceptance criterion 5: observable, maintainable, and integrable into existing workflows
The first four criteria govern "whether the answers are right"; the fifth governs "whether the system actually gets used and stands the test of time." Many Q&A systems perform flawlessly in demos but live permanently on a standalone test page: to ask a question, you first have to open that page, log in, and type. That form determines their fate: a month after launch, no one remembers where the entry point is, and the knowledge base stays frozen at the version from delivery day. To judge whether a system can go into production, first see whether it is willing to "disappear" into the places users already work.
The first consideration is the endpoint form factor. Employees spend their working day in chat windows and browsers, not at some newly opened URL. A deployable Q&A system should be embeddable as a web widget in an existing portal or documentation site, and mountable as a bot in DingTalk, Feishu (Lark), or WeCom, so users can ask questions and get answers directly within their existing conversations. Among open-source projects, PandaWiki, for example, supports all of these integration forms. This is not a nice-to-have; it pushes the "cost of use" close to zero, because users don't have to change any habits to ask a question. Conversely, if a system can only give you an isolated page, its real usage will almost certainly drop to zero, and even excellent retrieval quality will never translate into value.
The second consideration is how content gets in, and whether it can keep being updated once it's there. A knowledge base is not a water tank filled once; it is a pipeline that needs a constant, fresh flow. During acceptance, confirm whether the import channels cover your real content sources: can it ingest a web page URL directly, can it follow a site's Sitemap to pull in all of the site's documentation in bulk, can it subscribe to RSS so newly published content is ingested automatically, and can it handle local offline files? These channels correspond to different content lifecycles: website documentation relies on URLs and Sitemaps, dynamic news relies on RSS, and internal materials rely on file uploads. Missing one channel means a whole category of knowledge stays outside the system permanently, leaving a persistent blind spot in the answers. Whether the system can keep itself up to date is fundamentally determined by how complete these import methods are.
The third consideration is hidden in the architecture: invisible day to day, but everything when it comes to maintenance. For a system that must evolve over the long term, how the engineering is organized directly determines what it costs to change a single line of code. Two practices are worth confirming whether you are selecting a product or building your own. The first is how the LLM is loaded: the model client should be initialized only once as a singleton, with all subsequent requests reusing the same instance. Loading a model carries significant overhead, and rebuilding the connection for every request needlessly drives up resource usage and response latency; making it a singleton eliminates that waste outright. The second is module boundaries: split vector database reads and writes, text chunking and cleaning, model interaction, and upper-layer application logic into independent modules, each responsible for its own area. The payoff of this separation shows up in year two: when you need to swap vector databases, adjust the chunking strategy, or plug in another model provider, the changes stay confined to a single module instead of rippling through the whole system.
Taken together, these three points show that observability and maintainability are really two sides of the same thing: the system must let you see what it is doing (which pipeline was triggered, which retrieval came back empty, which module threw an error), and it must let you change one part of it without rewriting the whole. During acceptance, it's worth asking a simple question: could an engineer who takes over in six months safely replace one of the components without reading through the entire codebase? If the answer is yes, the system can truly be trusted to stand the test of time.
Comparing three deployment paths: custom-built RAG, open-source products, or low-code orchestration
The five acceptance criteria above define the boundaries of "usable," but deployment also involves an unavoidable decision: build it yourself starting from the vector database, or take an existing solution and adapt it? These three paths are not parallel options on a level plane; they are trade-offs across three dimensions: engineering investment, degree of control, and speed to launch. Lay out these three variables clearly, and the selection decision largely answers itself.
Start with building your own. The typical form of this path is a fully local stack: Milvus as the vector database, m3e (optimized for Chinese) as the embedding model, and a lightweight qwen2.5 as the generation model. Its core value comes down to one thing: data never leaves the building. The entire pipeline can run in an air-gapped environment with no API costs, and every parameter at every stage is in your own hands; similarity thresholds, retrieval strategies, and model weights can all be changed. The cost is just as direct: chunking strategy, embedding service, retrieval tuning, model deployment, and operations monitoring each have to be written and carried by your own team. Without stable R&D investment, a custom build can easily get stuck at the gap between "a demo that runs" and "a system that is usable." So this path makes sense for only one kind of team: one with hard data-privacy constraints that can also sustain an engineering team capable of continuous iteration. Teams that treat it as their first choice rather than a fallback often underestimate the long tail of maintenance costs.
The second path is to use an open-source product directly. These systems already package knowledge base construction, multi-source import, AI Q&A, and AI search into an out-of-the-box form that you can self-host with Docker. Take PandaWiki as an example: it has accumulated about 9.8k stars on GitHub and has shipped more than three hundred releases at a fairly high iteration cadence, so both its community activity and its maintenance continuity hold up to scrutiny. For teams without the time to build from scratch that still want their data on their own machines, this is the most cost-effective starting point. But keep an eye on two prerequisites. The first is the deployment bar: it requires a Docker environment of version 20.x or later on Linux, so the machines and operations need to be ready in advance. The second is licensing risk: it is licensed under AGPL-3.0, which means that once commercial distribution or externally provided services are involved, your modifications and integration code may be required to be open-sourced on the same terms. The implications of this for legal and your business model must be put on the table before you start, not discovered as unavoidable after launch.
The third path is low-code orchestration. It breaks the entire Q&A pipeline into visual nodes that you drag and connect on a canvas. A working multimodal Q&A flow is typically assembled from a handful of node types: user input, knowledge retrieval, conditional branching, a text generation model, and a vision model. The "retrieval first," "conditional routing," and "multimodal triggering" repeatedly emphasized in the acceptance criteria above map neatly onto a few visible nodes in this orchestration paradigm, each of which can be debugged on its own. Its biggest advantage is turning an abstract pipeline into an observable topology: if a step goes wrong, you open that node and look at its inputs and outputs; if you want to add a new branch, you connect two lines. For scenarios that are still validating business hypotheses and need rapid trial and error, it compresses the cycle "from idea to demo" to the minimum. The limitation is that deep customization runs into the ceiling of the platform's abstractions, and some fine-grained control is less direct than writing your own code.
Putting the three side by side in a table makes the differences clearer:
| Dimension | Custom-built RAG | Open-source product | Low-code orchestration |
|---|---|---|---|
| Speed to launch | Slow; built layer by layer | Medium; usable once deployed | Fast; assembled by drag-and-drop |
| Degree of control | Highest; full stack can be modified | Medium; constrained by the framework | Limited by platform abstractions |
| Engineering investment | Large, including long-term operations | Medium; mainly deployment and operations | Small; mostly configuration |
| Data ownership | Fully owned; can run offline | Self-hosted; data is yours | Depends on the chosen model and hosting method |
| Main constraints | R&D and maintenance capacity | Docker environment; AGPL obligations | Limited customization depth |
| Best fit | Hard privacy constraints + strong R&D | Self-hosting; quick start | Business validation; multimodal experiments |
One final point matters more than the selection itself: these three paths are not a mutually exclusive single choice. The pragmatic approach is to first use an open-source product or low-code orchestration to work through the five acceptance criteria above one by one, confirm that the Q&A system is genuinely "usable" on your own business data, and only then decide whether it is worth investing in a custom build for deep customization. Validating first and investing heavily later is far safer than betting on a custom build from the outset. The pitfall many teams fall into is precisely taking the heaviest path as their first step.
FAQ
The AI often makes up content that isn't in the knowledge base. How do we troubleshoot it?
This is the most common complaint after launch, but it is not a single problem; it is a class of symptoms. When troubleshooting, don't rush to tweak the prompt. First locate the point of failure by working backward along the pipeline.
Step one: look at the input the model received. Print out the retrieval results for a fabricated answer exactly as they were. Very often you'll find that retrieval returned nothing at all, or returned only irrelevant chunks. When a model has no usable context, it instinctively fills the gap with pretraining knowledge, and that is the direct source of hallucination. If this is the case, the problem is on the retrieval side, not the generation side, and changing the prompt won't help.
Step two: once you've confirmed retrieval is non-empty, check whether the prompt hard-codes the constraint "answer only based on the following content; if there is no basis, say plainly that nothing was found." If the constraint is missing or too soft (for example, only saying "please refer to"), the model will treat the retrieved content as a suggestion rather than a boundary. The wording here needs to be mandatory: it's better to have it answer "no relevant information was found in the knowledge base" than to give an answer that seems fluent but is actually fabricated.
Step three: check chunk granularity. If a passage is split in two with key information scattered across two chunks, and retrieval hits only one of them, the model will get half a fact and invent the other half on its own. In that case, adjusting the chunking strategy or increasing the overlap works faster than repeatedly rewriting the prompt.
A practical rule of thumb: ask the same question three times. If the fabricated content is different each time, retrieval is most likely not supplying enough context; if the fabricated content is consistent every time, it's more likely that a piece of wrong data has actually made it into the knowledge base, and you need to go back and clean the data source.
We uploaded an image, but the model says it can't recognize the file. What's misconfigured?
First separate two things: whether the image was stored, and whether the image was sent to a model that can understand it. If either step breaks, the symptom is the same ("unable to recognize"), but the fix is in a completely different place.
The most frequent cause is incorrect routing. Systems often have both a text-only model and a multimodal model attached. If, after a user uploads an image, the request is still dispatched to the text model by default, that model really can't see the image; it receives only a file path or garbled data, and so it replies that it can't recognize the file. To troubleshoot, check the request logs for which model endpoint was actually called. This step alone filters out most configuration problems.
The second category is format and size. Some models have hard limits on image format, resolution, and single-file size; anything over the limit is silently dropped or errors out during preprocessing. Start with a minimal test using a small, standard PNG or JPG image. If it works, the pipeline itself is fine, and you can go back to investigating the properties of the specific image.
The third category is mistaking "can store images" for "can read images." Some systems support uploading images as attachments to be saved and downloaded, but have no visual understanding capability connected; to the model, the image is just unreadable binary data. This isn't a configuration error but a missing capability, and you need to confirm that the chosen model itself supports visual input.
Can a knowledge base Q&A system be built in a fully offline intranet environment with no internet access?
Yes. In fact, for many teams with data compliance requirements, intranet deployment is the only acceptable option. But the costs need to be worked out in advance.
There are two core constraints. First, the model must run locally: you need to choose an open-source model that supports private deployment and provision the corresponding inference compute, which usually means at least one adequate GPU, with VRAM requirements rising as the model grows. Second, the embedding model must also run locally, because building the vector index and every retrieval depend on it; this step can't quietly call an external API.
Beyond the models, mature open-source options that can run offline exist for the vector database, document parsing, and application service layers, so assembling the stack is not hard. The real hidden cost is quality: at the same parameter scale, the answer quality of a small local model is usually below that of a large model called over the internet, and the gap is even more pronounced for capabilities such as multimodal understanding and complex reasoning. The pragmatic approach is to first define the minimum answer quality the business can accept, then work backward to how large a model you need, rather than assuming bigger is always better. For many internal knowledge Q&A scenarios, a mid-sized model combined with solid retrieval is enough.
Also plan for updates and maintenance: offline environments don't receive automatic updates, so model upgrades and security patches must all go through a manual process. Factor this operational workload into long-term costs rather than looking only at the one-time build.
We have no R&D team and a limited budget. What is the lowest-cost way to get started?
Don't start with a custom-built RAG stack. Building a retrieval-plus-generation pipeline from scratch, and just getting chunking, embedding, retrieval, reranking, and generation tuned to a usable level, requires sustained engineering time, which is exactly the resource a team without R&D lacks most.
A more sensible starting path is to use low-code orchestration or an off-the-shelf open-source Q&A product, and put your effort into data and acceptance rather than infrastructure. Concretely: start with a pilot in a scenario with clear boundaries and relatively well-organized documentation, such as product manual Q&A or internal policy lookup; the narrower the scope, the easier it is to produce noticeable results. Clean up that set of documents by removing outdated content, standardizing formats, and filling in missing heading levels. The return on this effort is far higher than fiddling with model parameters.
During validation, prefer pay-as-you-go online model APIs: get the workflow running and confirm the results first, then decide whether to switch to a self-built setup for compliance or cost reasons. This keeps upfront investment very low and avoids buying a pile of fixed assets before the value has been proven.
A counterintuitive but practical reminder: in the early stage, what deserves the most time is not product selection but preparing a list of test questions that covers typical phrasings. With this list, whichever solution you use, you can judge objectively whether it is actually usable, and you can quickly run regression checks when you swap tools later. That will save you more detours than any amount of agonizing over selection.