2026-07-20
Building an enterprise knowledge base: a complete plan from zero to launch
Building an enterprise knowledge base involves data collection, chunking, vector database selection, retrieval tuning, access control and more. This article lays out a practical path from zero to launch, covering data source prioritization, hybrid retrieval architecture design, phased rollout strategy and post-launch operations, helping teams avoid detours and quickly build an enterprise knowledge base that is usable, manageable and able to evolve.
Where the data comes from: inventorying sources and setting priorities
The most common way knowledge base projects fail is not a wrong technology choice but a wrong first batch of data: either too much low-quality content is poured in and retrieval results drown in noise, or high-frequency scenarios are left out and users are disappointed on their first try. Inventorying and prioritizing data sources is essentially answering an engineering question: which knowledge is worth paying to structure first?
Inventory in four tiers by data form
Enterprise knowledge is scattered across different systems, and differences in form determine integration cost and processing approach. We recommend taking stock tier by tier as follows:
| Tier | Typical carriers | Integration characteristics |
|---|---|---|
| Structured data | Business databases, CRM fields, ticketing systems | Clear schema and can be extracted directly with SQL, but requires masking and field concatenation |
| Semi-structured documents | PDF specifications, DOCX proposals, Markdown wikis | Varied formats and highly variable parsing quality; the parsing pipeline must be tuned format by format |
| Unstructured conversations | IM group chats, customer service conversations, meeting transcripts | Low signal-to-noise ratio; must be filtered, merged and summarized before it is worth ingesting |
| External real-time sources | Website pages, RSS feeds, third-party APIs | Content changes continuously; requires incremental crawling and version comparison |
Taking inventory isn't done once you've made a list. Each data source needs three attributes: where it currently lives, who owns the content and how often it is updated. This table is your knowledge map; every later decision about ingestion, updates and retirement is based on it. In practice, spending two to three weeks on a focused round of discovery to map the knowledge assets scattered across departments is far more effective than rushing to write code.
Use a priority matrix to decide the first batch to ingest
Don't try to pour all the data in at once. The criteria for the first batch can be judged with a simple two-dimensional matrix:
- Horizontal axis: query frequency. How many times a week is this type of question asked? Ticket statistics, search logs or customer service reviews all work as sources.
- Vertical axis: answer value. How much human intervention does answering this question correctly save? How much repeated back-and-forth does it avoid?
Content in the "high frequency × high value" quadrant is ingested first, typically product FAQs, standard operating procedures and pricing rules. Content in the "low frequency × low value" quadrant (such as archived historical files) is set aside for now. This screening keeps the first batch of data to a manageable size, preserving retrieval precision while letting pilot users see value quickly.
Engineering multi-source ingestion
Once the data scope is set, choose the ingestion method by scenario:
- Bulk file upload: suited to importing existing documents in bulk; supports common formats such as PDF, DOCX, TXT and Markdown. Watch how well the parser handles tables and nested lists.
- Web crawling and sitemap import: suited to content that already has an online structure, such as a website help center or product documentation site. A sitemap lets you get the full list of site URLs in one go and pull them in bulk.
- RSS feeds: suited to external sources that need continuous tracking, such as industry news and compliance announcements.
- IM message extraction: connect to the open APIs of platforms such as WeCom and Feishu, pull conversation records by channel or topic, and ingest them after summarization and cleaning.
Every ingestion channel needs a clearly designated knowledge owner: not the technical lead, but the business role responsible for the accuracy of the content. If a data source has no owner, it is better not to connect it for now. The core discipline of building a knowledge base is to focus on one or two high-value scenarios and close the loop first, then gradually expand data coverage. Skip inventory and prioritization and just pile on data, and you will most likely end up with something nobody uses.
How to split the data: chunking strategy and structuring
Once data sources have been inventoried, many teams feed documents straight into the vector database, treating chunking as just a technical parameter. This is the first pitfall. The chunking approach determines whether AI can "understand" the knowledge base rather than merely "store" it. AI can't understand knowledge that hasn't been structured, and this is the most underestimated part of building an enterprise knowledge base.
Structure first, then chunk
Before chunking, there must be a round of structural preprocessing; otherwise every chunk you cut is a fragment of text with no context. Four things need to be done:
- Heading hierarchy extraction: parse the document's first- and second-level headings into structured fields and store them as metadata with each chunk, so retrieval can apply secondary filtering by heading level rather than relying purely on semantic similarity to guess.
- Tables to key-value pairs: tables are where chunking goes wrong most often, and splitting them by row loses the mapping to column names. The right approach is to convert each row into key-value text of the form "field name: value," so that each fragment still expresses complete information on its own after splitting.
- Image OCR: policy documents and approval flowcharts often hide key information in screenshots. Without OCR, that knowledge simply disappears and will never be found by retrieval.
- Separate labeling for code blocks: code snippets in technical documents must not be split together with explanatory text. Label them separately to preserve complete code structure, so the chunking strategy doesn't break them into non-executable fragments.
How well this step is done directly sets the ceiling for chunking and retrieval later. The more complete the structuring, the more reliable the AI's answers.
Chunk differently by document type
Only after structuring is complete do you decide on chunk granularity. Here, one set of parameters can't fit everything:
- Technical documents (API docs, operating manuals, troubleshooting records): use small chunks, with chunk_size around 300 tokens. Technical content is semantically dense; a single step or parameter description is often an independent unit of knowledge, and only small chunks ensure that a retrieval hit returns the precisely matching passage rather than a large block mixed with lots of irrelevant context.
- Policy documents (HR policies, financial approval processes, compliance rules): use large chunks, with chunk_size at 800–1,000 tokens. Clauses in these documents often have strong logical dependencies. For example, "applicable conditions" and "exceptions" may sit in different paragraphs but must be understood together. Chunks that are too small lose the complete meaning, leading the AI to draw conclusions after seeing only half a sentence.
This differentiation is not optional; it is mandatory. When a knowledge base gives answers "taken out of context" after launch, the root cause is often a uniform chunk_size across all documents: technical documents cut too large produce noise, and policy documents cut too small lose context.
Overlap windows keep key information from being cut off
Whatever the document type, leave some overlap between chunks. The reason is simple: no fixed-length splitting rule will land exactly on semantic boundaries, and a complete causal explanation or a complete conditional statement may well fall right at the junction of two chunks. The overlap window lets the chunks on either side of a boundary each keep part of the adjacent content, so even if the split point isn't ideal, key information won't be chopped into two fragments that make no sense on their own. This parameter doesn't need to be tuned to perfection, but it must never be skipped. Without it, troubleshooting "the AI's answer misses the question" later becomes extremely difficult, because it is hard to tell whether retrieval went wrong or the chunking itself lost information.
What this step is really about
Chunking and structuring are not routine cleanup before ingestion; they are the primary determinant of knowledge base accuracy. "Garbage in, garbage out" applies right here: if chunk granularity is wrong, structuring is skipped and boundary information is cut off, then no matter how carefully the vector database is chosen or how finely the retrieval pipeline is tuned, you are optimizing on top of incomplete information, and the ceiling was locked in at this step. For critical business scenarios such as approval processes and compliance clauses, even when chunking and structuring are done well, we recommend keeping a human review step as a safety net. After all, a chunking strategy solves the problem of "information completeness"; it cannot replace the final check on business correctness.
Where to store the data: choosing a vector database and a hybrid architecture
Selection isn't a question of "which vector database is hottest"; it is a combined decision across four dimensions.
Data scale is the first dividing line. Up to around a million chunks, single-node deployments such as Milvus Standalone, Qdrant or even PostgreSQL + pgvector can handle the load, with low operations cost, and a team of one or two people can maintain them. Once you move past tens of millions toward hundreds of millions, you have to consider a distributed architecture: Milvus Cluster, Elasticsearch's vector search module, or a cloud provider's managed offering, because single-node memory and index rebuild time become hard bottlenecks. Estimate your real data volume before choosing, and don't use "we might need to scale in the future" as a reason to go straight to the heaviest architecture. Over-engineering early on only slows down the launch.
On deployment, core enterprise data carries privacy risks and can't be sent to third-party APIs, so the knowledge base needs to live in a private environment. This means vector databases holding sensitive data such as finance, contracts and customer information must sit in a data center or private cloud you control; managed services (such as Pinecone or cloud providers' vector databases) are suitable only for public materials or externally facing knowledge. This isn't a technical preference but a compliance red line. When selecting, first ask the business side "can this data leave the corporate network?" before looking at performance metrics.
Scalar filtering capability is often underestimated. The core flow of RAG is to retrieve relevant documents first and then have the LLM generate an answer, but the relevance of retrieval hits is only the first layer. In production, almost every query adds filter conditions: by department, by time range, by document permissions. If the vector database supports only pure vector retrieval, filtering has to be done as a second pass at the application layer, and both recall and performance suffer. Milvus, Qdrant and Weaviate all support filtering on scalar fields combined with vector retrieval. When selecting, be sure to measure retrieval latency with filters applied, rather than looking only at the pure vector search speeds in official benchmarks.
Community activity determines whether you can find answers when you run into problems. Projects with many Chinese contributors (Milvus) have richer Chinese-language materials and troubleshooting cases, while international projects (Qdrant, Weaviate) have high-quality documentation but weaker Chinese support. This dimension seems soft but actually affects operational efficiency later on; for problems such as index corruption and memory leaks in particular, community response speed directly determines recovery time.
A hybrid deployment architecture is the realistic choice for most medium and large enterprises: sensitive data and core business information are deployed on-premises, while general knowledge and public information can use cloud solutions, with the two connected through APIs. In practice, build a unified retrieval gateway that routes each query to the local or cloud store depending on the data domain involved, then ranks the returned results by relevance in a single pass before handing them to the LLM. This meets compliance requirements without incurring private deployment operations costs for public knowledge (such as product manuals and industry reports). The gateway layer needs sound timeout and degradation design, so that a cloud outage can't drag down the entire retrieval pipeline.
Metadata index design is easily overlooked but sets the ceiling on later capabilities. Every chunk should be bound to at least four kinds of metadata at ingestion: source document ID (for tracing and for locating content during updates), update time (to judge whether content is outdated), permission tags (mapped to org structure or roles, for pre-filtering at retrieval time) and business line (to support multi-tenant or department-level knowledge isolation). These fields are not after-the-fact patches; plan the write structure during the chunking stage. Otherwise, when you later implement access control or data cleanup, you'll have to tear down and rebuild the index, at a cost far higher than spending an extra half-day designing the schema at the start.
How to query the data: the retrieval pipeline and quality tuning
Whether a knowledge base is usable ultimately comes down to the retrieval pipeline. However neatly the data is stored, if retrieval fails, the answers users get will still go off track. This section breaks the retrieval pipeline into four engineering actions: hybrid retrieval, reranking trade-offs, query rewriting and evaluation.
Step 1: hybrid retrieval, keeping both semantics and keywords
Pure vector retrieval has a long-standing problem: with keyword-heavy queries involving model numbers, IDs or proper nouns, content that is semantically similar but doesn't match literally tends to crowd out the right results and send answers off track. The solution is a hybrid approach that layers BM25 keyword retrieval on top of vector semantic retrieval, with a recommended weighting of 0.7 for vector retrieval and 0.3 for keyword retrieval. The vector part understands roughly what the sentence is asking, while the BM25 part locks onto terms that must appear exactly. Results from the two paths are scored separately, fused and ranked by weight, then passed on for reranking. The weights aren't fixed once and for all; fine-tune them according to how keyword-dense the business content is. A legal contract library, for example, can reasonably raise the BM25 weight.
Step 2: coarse filtering first, then reranking; don't make the LLM read a pile of irrelevant documents
Retrieval shouldn't go straight to throwing 4 document chunks into the LLM in one step; leave room for reranking in between. In engineering practice this is usually done in two stages. The initial screening stage casts a wide net for candidates, so truly relevant chunks aren't missed because of errors in the initial ranking; then a Reranker model reranks this batch of candidates, rescoring them on semantic relevance, and only a small number of the most relevant results are sent to the LLM as context. This two-stage "coarse filter + rerank" design essentially uses a lighter but more precise model for a second pass, compensating for the rough scoring of the first-stage retriever. The reranking model can be an open-source cross-encoder or a cloud provider's Rerank API, depending on how the team trades off latency against cost.
Step 3: query rewriting, so the model doesn't have to take vague questions at face value
Users' original questions are often short, colloquial and unclear in their references, such as "how do I handle that reimbursement thing from last time," and retrieving on them directly yields a low hit rate. Add a query rewriting layer before retrieval. First recognize intent, determining whether the user is asking about a process, a rule or data; then expand the query, filling in synonyms, contextual references and likely technical terms to generate one or more query variants better suited to retrieval, retrieve on each, and merge the results. This step noticeably improves first-pass recall, especially for customer service and administrative knowledge bases, where user questions tend to be the least standardized. The rewriting logic can use a small model for lightweight intent classification plus template expansion; there is no need for a heavy model, which also keeps latency in check.
Step 4: build an evaluation set to quantify how good retrieval is
Tuning the retrieval pipeline can't rely on gut feel; it needs a repeatable evaluation mechanism. Build up a set of golden QA pairs, each containing a standard question, the documents that should be hit and a reference answer, covering high-frequency questions and edge cases. Run this evaluation set regularly and quantify three metrics: recall (did the documents that should be hit make it into the candidate set), precision (are the top results after reranking actually relevant) and answer relevance (how well the LLM's final answer matches the reference answer). Looking at the three metrics separately is critical. Low recall points to problems in hybrid retrieval or the rewriting strategy; low precision points to problems in the reranking model or the prompt; low answer relevance with the other two normal means the problem lies in the LLM's generation stage. Also be clear that the upper limit of AI Q&A accuracy depends on how well structured the knowledge base itself is and on the quality of its materials. No matter how finely the retrieval pipeline is tuned, if the underlying materials are messy and vague, answers will still be unreliable, so for critical business scenarios we recommend keeping a human review step as a safety net. The evaluation set should also evolve with the knowledge base: after each large batch of new documents or each change in chunking strategy, rerun the evaluation to confirm that production performance hasn't regressed.
Who can see what: permission models and security controls
Once a knowledge base goes live, the first question people ask is not "is retrieval accurate?" but "who can see this document?" If permissions are poorly designed, no matter how comprehensive the knowledge base is, nobody will dare connect it to production workflows. This is also a common reason knowledge base projects fail: companies won't use knowledge that lacks access control.
Document-level RBAC, with permissions carried in metadata. Don't "retrieve first, then filter" at the application layer; that wastes compute and risks returning content to unauthorized users. The right approach is to write permission tags into vector metadata at ingestion, with common dimensions including department, role, project team and document owner. At retrieval time, first generate a permission filter from the current user's identity (for example department in (...) or project_id in (...)), then run the vector search. In other words, "pre-filtering" rather than "post-filtering." Mainstream vector databases (Milvus, Weaviate, PGVector) all support conditional filtering at retrieval time with manageable performance overhead, so don't skip this step.
Classify sensitive data, and keep confidential content out of shared vector databases. We recommend tagging content at three levels, public, internal and confidential:
- Public: product documentation and FAQs, searchable by everyone and even eligible for cloud solutions;
- Internal: process standards and project materials, with visibility controlled by RBAC;
- Confidential: contracts, compensation, unpublished financial reports and the like. We recommend keeping these out of the unified vector database and using separate encrypted storage with an independent retrieval pipeline instead, or skipping vectorization altogether and allowing only restricted keyword-level queries.
This classification is not a one-time exercise. It should be labeled by the business side at ingestion, or screened first by rules plus a model and then reviewed by humans; otherwise it is easy to end up with a loophole where "highly sensitive documents are treated as internal documents."
Audit logs: leave a trail for every retrieval. Record at least four items: the identity of the querying user, the original query, the list of matched document IDs, and the final answer returned to the user. This isn't about "settling scores later"; it is a hard requirement for compliance review and permission audits. Many companies receive internal audit or regulatory inquiries after launching a knowledge base, and without logs they are essentially unable to prove anything. Store logs in a separate database, decoupled from the vector retrieval pipeline, so the audit system doesn't slow down the main retrieval path.
Private deployment keeps data inside the network perimeter. Core enterprise data carries privacy risks and can't be uploaded directly to third-party LLM APIs for vectorization or Q&A; this is the fundamental reason many companies choose private deployment for their knowledge bases. If full private deployment is too expensive, use a hybrid architecture: core business data and sensitive materials are deployed on-premises (a local vector database plus a local or privately deployed model), while public materials and general knowledge use cloud solutions. The two are connected through APIs, with the retrieval layer orchestrating centrally and routing to different execution environments by sensitivity level. This controls cost while confining the exposure to data leaks to the public tier.
To sum up this section's engineering decisions in one sentence: write permission tags into metadata, filter before retrieval, physically isolate sensitive data, log the entire chain, and keep core data off the public network. Only with all five in place does a knowledge base have the prerequisites for being genuinely trusted and used by the business.
How to launch: a phased path from pilot to full rollout
The most common way knowledge base projects die is not a wrong technology choice, but a one-shot full launch that gets buried under an avalanche of bad cases, leaving the team firefighting and user confidence at zero. The right rhythm is to break the risk into small pieces: close the loop on one business line first, then expand gradually.
Which business line to pilot
There are only two criteria for choosing: the best data quality and the highest usage frequency. Good data quality means you don't have to solve "content governance" and "system tuning" at the same time during the pilot; high usage frequency means you can accumulate enough real query samples in a short period to expose problems. Typical preferred scenarios include product FAQs, internal IT operations manuals and standardized business process documents. This content is highly structured, with clear answer boundaries, making it easy to judge right from wrong.
Keep the pilot to a few weeks. Complete data ingestion and basic integration testing first, then open it to a small group of seed users for intensive use, and afterward adjust the chunking strategy, tune retrieval parameters and fill in missing documents based on feedback. If core metrics still fall short after four weeks, the problem most likely lies in the data sources themselves rather than the system; at that point, go back and fix the fundamentals rather than forcing a launch.
How to judge acceptance
The end of the pilot needs a set of quantifiable pass criteria, with the team aligning on expectations in advance so the launch decision doesn't become a subjective tug-of-war. There are three core dimensions:
| Dimension | Focus | Assessment method |
|---|---|---|
| Answer quality | Whether returned content is accurate, complete and free of hallucination | Human spot checks + automated comparison against a golden QA set |
| Response efficiency | Whether end-to-end latency is within a range users find acceptable | P95 latency monitoring |
| User experience | Whether actual users are willing to keep using it | Short survey or NPS collection |
Specific thresholds should be agreed jointly by the business and technical sides, because tolerance for inaccuracy varies widely across scenarios (compliance scenarios are far stricter than everyday administrative Q&A). The key is to fix the numbers in writing in advance, so the parties don't apply different standards at launch.
How to ramp up during the phased rollout
Once the pilot passes acceptance, don't push it to everyone at once. We recommend three steps:
- Small-scale internal beta: have a small group of users in the target department use it first for one to two weeks, focusing on collecting bad cases and fixing them one by one. The goal at this stage isn't coverage but driving down high-frequency errors.
- Full department: once a high share of the bad cases from the internal beta has been fixed, open it to the whole department. As user numbers grow, long-tail problems and concurrency bottlenecks will surface, and operations needs to watch queue backlogs and cache hit rates.
- Cross-department rollout: after a single department has run stably for two to three weeks, replicate to the next department. Each time you add a department, reassess that department's data readiness. Directly reusing the previous department's configuration often backfires, because document styles, terminology and permission structures all differ.
Human fallback mechanism
For critical business scenarios (interpreting contract clauses, compliance policy inquiries, queries on sensitive customer information and so on), the system should not give final answers on its own. The engineering approach is to have the model output a confidence score along with its result and, when confidence falls below a preset threshold, automatically route the request to a human review queue. The user sees "forwarded to a specialist for confirmation" rather than a potentially wrong answer.
The value of this mechanism lies not only in safeguarding accuracy but, more importantly, in building user trust. In the early stages, users' trust in AI answers has to accumulate gradually, and one serious error can lead an entire department to abandon the system. The share of requests needing human review will fall as the model is tuned and the knowledge base matures, but we recommend reserving enough staffing budget in the early days after launch.
Common failure points
- Skipping the pilot and going straight to full rollout: data problems are magnified in front of every user, and fixes can't keep up with complaints.
- Piloting the hardest scenario: in trying to prove technical capability, the team exhausts the timeline on data governance and never produces positive feedback.
- Not collecting feedback during the phased rollout: traffic is ramped up without a channel for collecting bad cases, and by the time problems pile up too high to ignore, the best window for fixing them has passed.
- No exit mechanism: if core metrics keep deteriorating during the phased rollout, there needs to be a predefined rollback plan rather than stubbornly pushing on.
The essence of a phased rollout is trading controlled risk for real feedback. Each ramp-up is an experiment to test a hypothesis, not an administrative approval checkpoint. Keep that mindset, and the launch won't turn into a gamble.
How to keep it alive: ongoing operations and keeping data fresh
Launching the knowledge base is only the starting point. What really decides success or failure is whether, within three months of launch, a closed loop forms: "data comes in → gets used → feedback comes back → quality goes up." Most enterprise knowledge base projects stall not because of wrong technology choices but because nobody keeps feeding them.
Incremental sync: don't let full rebuilds drag you down
Re-vectorizing everything after each document change is basically unacceptable once a knowledge base exceeds tens of thousands of documents. The more sensible engineering approach is an event-driven incremental pipeline:
- Document systems (Confluence, Feishu Docs, SharePoint) detect changes via webhooks or scheduled polling and push change events to a message queue
- Consumers receive the changed document and re-run chunking and vectorization for that document only
- Document ID + version number serve as the linking key for vector records; after a new version is written, all vectors belonging to the old version are deleted
- For partial updates (such as a change to one paragraph in a document), run a paragraph-level diff and rebuild only the affected chunks
Key design principle: vector records must retain enough metadata (source document ID, version number, chunk position index); otherwise, during incremental replacement you can't precisely locate the old vectors that need to be retired.
Expiry cleanup: give knowledge a shelf life
Technical documents, product manuals and policies are all time-sensitive. Without expiry governance, outdated information creeps into retrieval results and user trust drops quickly.
- At ingestion, tag each piece of knowledge with a validity period field (set manually by the knowledge owner, or with defaults by document type, for example a default validity period for meeting minutes)
- At retrieval, down-weight expired documents (multiply by a decay factor) rather than deleting them outright; some historical information still has reference value
- Run expiry scans regularly, mark documents that have expired and not been renewed as "archived," and remove them from the main retrieval index
Knowledge owner accountability: make sure someone looks after the content
Enterprise knowledge is scattered across different teams, and knowledge domains without a clear owner eventually rot. Recommendations for putting this into practice:
- Divide knowledge into blocks by business domain, and assign a small number of owners to each block, responsible for content updates and regular inspections
- An owner's core responsibility is not "writing documents" but deciding which knowledge should go into the knowledge base, which should be retired and which needs to be added
- In step with the regular inspection cadence, owners spot-check retrieval hits in their knowledge domains and flag inaccurate content
Teams that have done a "knowledge map" survey will find this step much easier. Mapping where each type of knowledge is stored and who is responsible for it early on essentially lays the groundwork for the owner system.
User feedback loop: turn complaints into fuel
Every "thumbs down" from a user is free labeled data. On the engineering side, make this path work end to end:
- Provide a lightweight feedback entry point on the front end (thumbs up/down, with an optional reason)
- Thumbs-down records automatically enter a bad case queue and are dispatched to the relevant owner by knowledge domain
- The owner determines the root cause: is context missing because chunk boundaries were badly placed? Is the knowledge itself outdated? Or is this knowledge missing altogether?
- After a targeted fix, mark the case as resolved and add it to the regression test set
Once this mechanism is running, knowledge base accuracy keeps improving as usage grows: the more it is used, the more feedback it gets, and the higher the quality. What matters is not how many documents are stored, but whether knowledge can be retrieved accurately and kept usable over time.
FAQ
Is it worth building an enterprise knowledge base for a team of fewer than 50 people?
Yes, but in a different way. Small teams' knowledge management pain points are often more concentrated: new hires ramp up slowly, departing veterans take their experience with them, and the same questions get asked over and over. We recommend starting with one high-frequency scenario (such as common customer questions or an internal technical FAQ) and using a lightweight solution to validate value quickly. There is no need to ingest all knowledge from the start; first get the retrieval pipeline working with a dozen or so core documents, confirm that it really saves time, then expand gradually.
Should the vector database be privately deployed or cloud-hosted?
It depends on data sensitivity and operations capability. For scenarios involving core business data and customers' private information, private deployment is a hard requirement; a company can't hand sensitive materials to a third party to host. If the knowledge base mainly contains public product documentation or general technical materials, a cloud-hosted solution can greatly reduce the operations burden. The middle ground is a hybrid architecture: sensitive knowledge domains use an on-premises vector database, general knowledge domains use cloud services, and retrieval is routed to different backends according to permissions.
What if answer accuracy is low after the knowledge base goes live?
First determine whether the bottleneck is in retrieval or in generation. The method is simple: pull a batch of bad cases and manually check whether the top-K chunks returned by retrieval contain the correct answer. If the chunks contain the correct information but the final answer is wrong, the problem lies in the prompt or the model's comprehension; if the chunks didn't hit any relevant content at all, the problem lies in the chunking strategy or the retrieval configuration. Common retrieval-side fixes include adjusting chunk granularity, adding synonym mappings and adding metadata filter conditions. On the generation side, focus on how context is organized in the prompt and on instruction constraints.
How do you measure the ROI of a knowledge base?
Avoid vanity metrics such as "number of documents" or "number of calls." We recommend tracking three types of efficiency data that can be clearly attributed:
- Reduction in repeat inquiries: compare how often similar questions are asked in internal IM groups or ticketing systems before and after launch
- New-hire ramp-up time: whether the number of days from joining to handling work independently has shortened
- Time spent finding knowledge: through sample surveys or event tracking, compare the time it takes to "dig through documents for answers" with getting them "directly from the knowledge base"
These metrics don't need to be precise to the decimal point; a visible trend of improvement is enough to justify the investment. The key is to set baselines before launch; otherwise, data gathered after the fact will hardly be convincing.