Teverant AI · Insights

2026-06-25

Before you feed it to AI: an engineering checklist for enterprise document governance

When RAG performs poorly, the root cause is often not the model but the quality of the documents before they are ingested. Covering content cleaning, structured parsing, freshness management, permission syncing, and compliance boundaries, this article lays out a complete engineering checklist for enterprise document governance, helping teams build a solid data foundation before they build an AI knowledge base and avoid the systemic risk of garbage in, garbage out.

When RAG fails, the root cause is often before ingestion, where you can't see it

When a RAG system underperforms after launch, a team's first reaction is usually to suspect that the model isn't strong enough or that the wrong retrieval algorithm was chosen. So they start swapping embedding models, tuning similarity thresholds, and adding reranking, only to see little improvement after weeks of effort. This line of troubleshooting often goes astray because it assumes something by default: that the content going into the vector database is clean, accurate, and meant to be seen. But in most of the cases I've worked on, the problem lies exactly where that assumption doesn't yet hold: the step before ingestion was skipped.

Breaking it down makes this clearer. Retrieval and generation sit at the end of the pipeline, and they faithfully process whatever they are fed. If what goes into the vector database is documents full of formatting junk, version chaos, and missing permission labels, then no matter how precise the downstream stages are, the output is just "a high-quality restatement of the wrong material." In other words, the model has no way to judge for you whether a document is an old policy that was retired three years ago, nor can it tell whether a given passage should be visible only to the finance department. These judgments have to be made before the data enters the database; miss that window and the contamination gets baked into the index.

Outdated content is the most insidious category. Unlike garbled text, it isn't obvious at a glance; it looks exactly like a valid document and may even score very high on semantic matching. Picture an expense policy from two years ago: clearly itemized and properly worded. At retrieval time it will be reliably pulled back, and the answer the model generates from it will look professional and confident, while in reality it is guiding today's decisions with a set of rules that has already been abolished. What makes these errors dangerous is that they don't throw errors; users have a hard time noticing that the answer they received is out of date. Over time, valid content and zombie documents pile up together in the knowledge base, and it gradually turns into a "document graveyard" that nobody dares to trust and nobody wants to clean up. Content governance without clear ownership tends to drift out of control, which is exactly why, later on, I'll argue that every document should be tied to an owner and a review cycle, and that expired documents should either be taken down or explicitly flagged as questionable in search results.

Another category of problem is even more serious, because it involves security. When building an index, many teams care only about "what the content is" and forget to sync "who can see it." When a document's access control information doesn't follow it into the retrieval layer, the vector database becomes one big pool that treats every querier the same. At that point the risk is no longer wrong answers but unauthorized access. An ordinary employee who should never have access to compensation data may not need any hacking at all; simply rephrasing a question as a semantic search can get the system to return restricted content to them as if it were an ordinary knowledge snippet. The permission boundaries set up in the original document system are bypassed through the retrieval channel. This is permission leakage.

There's an easy misconception here: thinking it's enough to filter by permissions once at ingestion. Static, snapshot-based permission handling can't withstand reality. Org structures change, people move around, project permissions are adjusted all the time, and someone who has access today may transfer to another role next month. If the retrieval layer relies on the permission state frozen at the moment of ingestion, sooner or later it will fall out of step with actual authorization. The right approach is to perform the permission check at the moment the query happens: when a user runs a search, filter the candidate results in real time using that user's current identity and authorizations, rather than trusting labels written into the index months ago. I'll go into the engineering implementation in the section on permission syncing.

Put these three hazards together (dirty data, outdated content, and mismatched permissions) and you'll see they share one trait: they all happen before or at the moment data enters the database, and once they happen they're very hard to fix downstream. You can't turn a retired document back into a valid one by tuning retrieval parameters, and you can't fill in missing permission labels by switching models. So the next few sections of this series focus not on models and algorithms but on the often-overlooked steps before anything is fed to AI: how to clean content and scan it for security issues, how to turn unstructured material into structured data that can be governed, how to manage freshness, how to keep permissions aligned in real time, and, once AI agents start writing back into the knowledge base, how audit and tiering provide the backstop. Only when you hold the line before ingestion does it make sense to talk about retrieval quality at all.

Step 1: content cleaning and security scanning before ingestion

Before you feed documents into a retrieval system, you have to admit one thing: most enterprise documents weren't prepared for machines. They were written by people for people, and they contain keys pasted in casually during debugging, customer phone numbers copied into support tickets, and temporary tokens left behind by an engineer on an internal wiki page. In the original documents these are nearly harmless, because accessing them requires permission to access that document in the first place. But once they enter the vector index, the boundary breaks: a configuration snippet that only the ops team could see may now be retrieved because it's semantically similar and stitched into an answer shown to anyone who runs a query. The essence of the problem isn't that the documents are dirty; it's that retrieval strips documents out of their original access context.

So the first class of problem the cleaning step has to solve is identifying and handling sensitive information. There's a timing decision to make explicit here: the detection point must sit at the ingestion stage, not in a scan after the fact. Remediating after ingestion means the sensitive data has already been chunked, vectorized, and written into the index, so what you have to clean up isn't just the original text but every derived embedding and all the metadata, and at any moment before you finish, it could be hit by a search. The benefit of moving detection upstream is that it stops the risk before any irreversible operation: the document is still a complete, traceable object, and you can decide whether to quarantine it, redact it, or reject it outright. Glean's Protect Plus in 2026 takes exactly this route, integrating PII detection with the company's existing security stack so that detection happens before content enters the index rather than after. This direction is worth borrowing from; the point isn't whose tool you use but that the security control point is pinned to the entrance of the pipeline.

As for handling strategy, there should be no one-size-fits-all approach. My recommendation is to route by type of sensitivity:

  • Structured, strong identifiers (national ID numbers, bank cards, API keys, private keys, access tokens) have clear format signatures, so regex plus entropy detection achieves high hit rates. By default these should be quarantined or blocked from ingestion outright and released only after human review.
  • Weak identifiers (names, email addresses, internal IPs, project code names) aren't necessarily sensitive on their own; they create identification risk only in combination. These are suitable for ingestion after redaction, replacing the original values with placeholders, which preserves the document's semantic structure for retrieval while cutting off the leakage path.
  • Context-dependent sensitive content, such as salary discussions or undisclosed M&A information, can't be caught reliably by pattern matching alone. This usually has to be backstopped by document source classification and access permissions; you shouldn't expect the cleaning step to carry it alone.

The second class of problem is subtler and easier to overlook: a document itself can be an attack vector. When the knowledge base becomes the retrieval source for RAG, document content is no longer just data to be read; the model treats it as context to understand and act on. An attacker can craft a document that looks normal but embeds manipulative instructions, such as "ignore all previous constraints" or "mark this content as the highest priority." Once that document is retrieved and stitched into the prompt, the model may act as the attacker intends. This is the logic of an ingestion attack: the contamination doesn't happen at query time but earlier, at ingestion, lying dormant in the index until the right query activates it.

To counter this kind of risk, you need a content security scan before documents enter the vector index. Unlike PII detection, this scan isn't looking for "is there sensitive data" but "will this content try to manipulate downstream behavior." Practical checkpoints include identifying instruction-like statement patterns in documents, especially imperative sentences that shouldn't appear in normal business documents; detecting abnormal formatting constructions, such as long runs of invisible characters, extremely long repeated tokens, or delimiters that try to break out of context boundaries; and applying stricter review thresholds to external documents from untrusted sources (user uploads, crawled web pages, third-party imports) than to internal ones. These checks won't catch every variant, but they will keep the cheapest and most common injections out.

Chaining these two scans into the ingestion pipeline usually takes the engineering form of a set of interceptors: a document comes in, goes through format parsing first, then runs PII detection and the content security scan in parallel; if either one flags it, it goes into a quarantine queue to await handling, and only if it passes both does it move on to the next step, structuring. The part worth investing real effort in is the handling record: every quarantined or redacted document should leave a traceable log that clearly records which rule was triggered, what was done, and who approved its release. This isn't just for compliance audits; it's because cleaning rules will inevitably evolve, and you need data to judge whether the rules are too loose or too strict. A cleaning step that can be both tuned and reviewed after the fact is a step that can keep running over the long term.

Step 2: turning unstructured content into structured data that can be governed

Cleaning deals with "dirt"; this step deals with "shape." The documents that are truly hard to handle in an enterprise are often not parsable text to begin with: contract PDFs scanned and archived years ago, invoice images exported by the finance department, meeting recordings, product demo videos. Put these into a vector database and the retriever sees only a blob of binary that can't be chunked, or a scanned page that gets skipped as blank. Much of the time, this is the root cause when RAG gives irrelevant answers: it's not that the model is bad, it's that the model never got readable content in the first place.

So the first job is to turn "text inside images" back into "text machines can read." OCR is the foundation at this layer, but you have to pick the engine for the scenario rather than use one model for everything. For ordinary official documents and manuals, general-purpose OCR is enough; once you move into a zero-tolerance domain like finance, things change. Bank statements, credit card bills, and VAT invoices are tables full of digits and amounts, and if a decimal point is off by one place or a thousands separator is read as a period, downstream reconciliation falls apart completely. These scenarios need recognition engines trained specifically on financial tables, optimized for digits, amount fields, and table lines, which can output statements directly as rows and columns in Excel or CSV rather than a string of scattered characters that has lost its positional relationships. Vendors specializing in financial document digitization typically advertise digit recognition accuracy on the order of 99%. What matters about this metric isn't its absolute value but that it shrinks the human review workload down to an acceptable spot-check range, and that is the dividing line for whether something can go into production.

But turning images into text is only the starting point of structuring, far from the end. What really determines whether a document can be "governed" is whether it carries a set of fields that can be searched, filtered, and audited. The body text of a contract alone isn't enough; the system needs to know whether it's a contract or an invoice, who the signing parties are, the effective and expiration dates, the amount involved, and which department it belongs to. That's metadata. Entering it by hand simply isn't realistic for tens of thousands of documents, and the quality of manual entry varies widely. A workable approach is to hook an automatic extraction step into the ingestion pipeline, so that after the model reads the content, it fills in these fields along the way: document types classified automatically, key entities recognized automatically, tags applied automatically. That way, every document arrives in the database carrying its own "ID card," and whether you later isolate permissions by department or trigger freshness alerts by expiration date, there are fields to rely on.

None of this is cutting-edge anymore; mature content management platforms have already made it standard. A fairly typical combination of capabilities is using AI to automatically classify and intelligently tag documents, completing metadata extraction and content analysis at the moment of ingestion, and then using that to provide more accurate search suggestions. In other words, classification and tagging are no longer manual work for archivists but a step that happens automatically as content comes in the door. The engineering value behind this is that it turns "governance" from an after-the-fact action requiring dedicated staff into an inline step in the ingestion chain: however many documents you have, governance automatically keeps up, and it doesn't break down just because the volume is large.

Among heterogeneous content, audio and video are the easiest to miss. A lot of a company's real knowledge (an architecture review, a customer Q&A session, a product training) sits in cloud drives as audio and video recordings and has never entered any searchable system. The key action for handling it is transcription: converting speech into timestamped text, then running it through the same cleaning, extraction, and tagging process as ordinary documents. Some platforms have already bundled transcription, image content recognition, and key information extraction into a single skill framework, which means you don't have to build a separate pipeline for each media type; audio goes in and indexable text passages come out. For engineering teams, this means a single, unified ingestion standard can absorb text, tables, images, audio, and video, sources that originally came in very different formats.

Looking at these layers together, the goal of this step is actually quite clear: whatever form content starts in, everything that enters the database should ultimately converge on the same standard format with fields. A comparison table helps show which processing steps each type of content needs to go through:

Original formCore processingForm after ingestion
Scans / image PDFsGeneral-purpose OCR + layout reconstructionPlain-text blocks that can be chunked
Financial receipts / statement tablesFinance-specific OCR that preserves row and column structureStructured tables (Excel/CSV)
Contracts, reports, and other long documentsMetadata extraction + automatic classification and taggingText with fields such as type, parties, and dates
Audio / videoSpeech transcription + subsequent extraction and taggingTimestamped, searchable text

One caveat: neither OCR nor model extraction is 100% reliable, and automated results for high-risk fields in areas like finance and legal shouldn't go straight into production. The pragmatic approach is to attach a confidence score to extraction results: anything below the threshold goes into a human review queue, anything above it is released automatically, and review records are kept as an audit trail. That way you get the throughput of automation while keeping a line where people can backstop the critical fields. Once structuring reaches this level, permission isolation and freshness management have something to grip. The next section picks up from here: after a document enters the database, how do you assign it an owner and a shelf life?

Step 3: freshness governance, giving every document an owner and a shelf life

The first two steps solve the problem of "dirty"; this one solves the problem of "old." A cleanly formatted, well-structured document whose content has expired is more damaging to a retrieval system than a block of garbled text. The vector model will likely treat garbled text as noise and push it to the margins, but outdated documents tend to be well written, confidently phrased, and semantically clear, which makes them exactly the ones most likely to be retrieved at the top of the results. A user asks "Which process do expense approvals go through?" and the system retrieves a carefully worded old policy; the model generates an answer from it that looks completely credible, and nobody questions it. This is where the problem quietly occurs: the model has no sense of time, and it doesn't know this document is a version that was retired last year.

So the first thing in freshness governance is to bind two mandatory fields to every ingested document: an owner and a review cycle. The owner determines whom the system should remind when the content needs updating; the review cycle determines when the document must be reconfirmed as valid. These fields sound like management actions, but in engineering terms they are metadata constraints: if either is missing when a document enters the knowledge base, the ingestion process should stop it rather than let it through and hope someone fills them in later. Once it's let through, you'll find six months later that nobody remembers who's responsible for the document. It becomes an orphan that will be neither updated nor deleted, lying quietly in the index waiting to mislead the next person who asks a question. That's how "document graveyards" form: not because there are too many documents, but because nobody is responsible for the validity of any individual one.

The second thing is to make the "expired" state visible in search results. Many teams simply delete documents when they expire, which is actually too crude. A policy may be entirely retired, or only one clause may have changed; deleting it outright also throws away its context and historical basis. A safer approach is to attach a freshness marker to documents and, when a search hits expired content, explicitly flag a warning in the results instead of silently returning it. The warning should be shown to the end user and also passed into the context sent to the model: when assembling the prompt, you can inject "This document ceased to be valid on such-and-such date" as a piece of metadata, giving the model a chance to avoid it or proactively point it out while generating. Only when the retrieval layer is aware of freshness will the model stop answering earnestly from a scrapped draft.

The third thing is retaining and tracing versions. Enterprise documents are rarely finalized after being written once; a policy may have gone through a dozen or more versions, while the retrieval system should only ever hit the version currently in effect. This requires the knowledge base to achieve two things at the underlying level simultaneously: "historical versions are available for lookup" and "the current version is unique." You need to be able to go back to any historical version to see what it said at the time, while ensuring that what the vector index and retrieval channel expose to the model is always the version in effect. These two goals are easily set against each other: dump every version into the index so you can trace history, and the same question retrieves three contradictory versions; keep only the latest version for cleanliness, and when a dispute arises you can't find the original basis. The right way to split it is to separate storage from retrieval: keep the full history in the document repository, attach only the current effective version to the retrieval index, and update the index pointer in sync whenever the version changes.

Put these three things together and you'll see that freshness governance essentially gives static documents a lifecycle: bind an owner at creation, review on a cycle while in use, flag rather than silently serve after expiration, and switch the effective version while retaining old ones on update. If any link in this chain is missing, the knowledge base's credibility will decay over time. This is also why many RAG systems perform well right after launch but start making frequent mistakes after six months: nobody maintains the content, the share of old documents in the index keeps rising, and retrieval quality naturally slides.

One more thing must be made clear: there's no one-size-fits-all model for freshness governance; its complexity scales directly with the size of the organization. For a five-person team, the content lifecycle may not need to be systematized at all: whoever wrote something remembers it, when it expires someone mentions it and it gets fixed, and keeping the review cycle in your head is good enough. But for a company of two thousand people, the situation is completely different. Documents span departments and business lines, authors may have left long ago, and the retirement of one policy can ripple through multiple downstream processes. At that point, owner, review cycle, and version status must be fields the system strictly enforces, because people simply can't remember them all. So don't copy a big company's governance framework onto a small team, and don't manage a large-scale knowledge base with a small team's casual mindset. First figure out what scale you're operating at, then decide how heavy your freshness governance needs to be. On this issue, over-engineering and missing governance both cause problems.

The test is actually quite simple: open your knowledge base, pick ten documents at random, and see how many you can immediately say who owns, when they were last reviewed, and whether what's currently in the index is the effective version. If you can answer, freshness governance is working. If you can't, what you're feeding the AI isn't knowledge but a pile of old files nobody vouches for.

Permission syncing: retrieval must check permissions at query time, not at ingestion

Many teams treat permissions as a one-time action during ingestion: tag the documents as they enter the database and assume that subsequent retrieval will naturally be controlled. This is a dangerous misunderstanding. A vector database is essentially an index of semantic similarity; it remembers "which passage is closest to this question," not "who is allowed to see this passage." If you don't check permissions again at query time, semantic search becomes a back door around your existing access control. An engineer without HR permissions could entirely plausibly fish out content locked in the HR directory with a question like "What was the company's average raise last year?" He can't see the file in the document list, but RAG will stitch sentences from it into the answer. This kind of "permission leakage" isn't a configuration mistake; it's a step missing at the architecture level.

So the right order is to think in reverse: first determine "who is asking this question and what can they see," then decide "which vector chunks are allowed to take part in retrieval." Permission filtering must happen at the retrieval stage, not the generation stage. If you wait until the model has already read unauthorized content into its context and then rely on the prompt to make it "pretend it didn't see it," you're handing the control point to a probabilistic system. There's some probability it will fail, and compliance audits don't accept probabilities. A workable engineering approach is to have every vector chunk carry access control metadata (owning directory, classification level, visible roles/groups), and at retrieval time push the current user's effective permissions down into the vector database query as a hard filter, removing unauthorized chunks first and then ranking by similarity. That way unauthorized content never even makes it into the candidate set, and the model has nothing to leak.

For this kind of real-time check to work, identity has to be trustworthy and unified. That means the retrieval service can't maintain its own user table; it has to connect to the company's existing identity source, getting the login state via SSO and resolving the user's current groups and roles through a directory service (LDAP/AD or an IdP). The key word is "current": transfers, project endings, and permission revocations must take effect on the very next query, not wait until the vector index is rebuilt to sync. Baking a permission snapshot into a static index leaves people who have left or changed roles with a window of unauthorized access. Having each query consult the live identity system adds an extra hop of latency every time, but it's the baseline cost of compliance.

The foundation is that documents themselves must first have clear ownership and classification. If all material is laid out flat in one big pool, with no categories and no read/write boundaries, there's no dimension to filter on at query time. A sensible way to organize things is to build a category tree by business domain or team and configure read and write permissions separately at both the category and document levels, so that "who can retrieve it" and "who can write revisions" become attributes that can be declared, inherited, and audited. Some document management platforms have already combined this with AI capabilities: automatically identifying document types, scanning for sensitive information in them (national ID numbers, contract amounts, customer lists), and using that to suggest classification levels or even tighten permissions automatically. This kind of automated detection can't replace human classification, but it can compress the needle-in-a-haystack problem of "which documents might be misclassified" into a reviewable list of candidates, significantly reducing the blind spots in manual governance.

Finally, there's audit. Once permission-aware retrieval goes live, the question you need to answer is no longer just "is the system secure?" but "at 3 p.m. last Tuesday, what did so-and-so search for, which documents were hit, and did any of them touch highly classified content?" For every query, the requester, the chunks that were hit, and the filter rules applied should all be written to an audit log, and the log itself must be tamper-proof and traceable. This is both the basis for assigning responsibility after an incident and, in normal times, a signal source for spotting abnormal access patterns (such as one account repeatedly probing sensitive directories within a short time).

Stack these layers together and the gap between a small team's shared drive and an enterprise knowledge base becomes clear. Five people sharing one space can keep order through mutual understanding; two thousand people, dozens of business lines, plus compliance reviews require access control, identity integration, lifecycle management, and audit to be built as system capabilities. Permissions aren't an add-on feature of a knowledge base; they determine whether the system can enter a real enterprise production environment at all. If it can't filter in real time at query time, RAG can only stay at the demo stage.

Data sovereignty and compliance boundaries: what to confirm before document chunks leave your walls

The previous steps addressed "whether the documents are clean"; this section addresses "where the documents go." Here is a fact that's easy to overlook: when you feed a contract, a medical record, or an employee file into a RAG system, it usually doesn't stay on your servers. The first step of retrieval augmentation is chunking the text and calling an embedding API to compute vectors; the last step of question answering is stitching the retrieved chunks into a prompt and sending it to a large language model (LLM). For a significant share of commercial knowledge base products, both steps go through third-party cloud APIs. In other words, passages of your original text actually leave your perimeter.

For most internal knowledge bases, this isn't a big deal. But for certain categories of data, leaving the country is itself a legal event, not a technology choice. EU personal data governed by the GDPR, specific data that requires a cross-border security assessment before export, and sensitive information defined by industry regulators: once chunks of this data enter an overseas embedding service or LLM inference cluster, you may have triggered a cross-border transfer without knowing it. In an audit, it won't show up as "called an API"; it will show up as "personal data flowed to a processor not on the compliance list."

So the questions to ask during selection aren't "how good is the model" but three harder ones: can embedding be self-hosted, can inference run on-premises or in a designated region, and will the vendor give written guarantees about data residency? If any of these three can't be answered, rule the option out for data subject to cross-border restrictions. It's worth pointing out that SSL encryption in transit and security certifications on the storage side address "security in transit and at rest," which is a different matter from "which jurisdiction the data is processed in." However complete the former is, it doesn't answer the cross-border question for you. For scenarios with strict data sovereignty requirements, a solution that supports private deployment and keeps the entire chain inside your own data center or your own VPC is often the only option that can pass a compliance review.

Turning these judgments into an actionable checklist is the safer route:

  • Classify first: tag which document repositories contain personal data and which contain export-restricted data before ingestion, rather than digging through them after something goes wrong.
  • Map the data flow: draw the complete path each category of data takes, from chunking and vectorization to inference, clearly marking the processor and physical location of every hop.
  • Enforce regional constraints: require self-hosted embedding and on-premises or designated-region inference for restricted data, enforced at the configuration level rather than relying on people to remember.
  • Backstop with contracts: for the parts that must use external services, make sure the data processing agreement, residency commitments, and retention and deletion clauses are all spelled out in black and white.

The other half of compliance is "being able to explain it." Regulators and auditors are increasingly unsatisfied with verbal assurances like "we're very secure"; they want to see whether you can reconstruct how the entire system was trained, monitored, and evaluated. That's exactly what frameworks such as the EU AI Act, the NIST risk management framework (NIST RMF), and ISO/IEC 42001 are after. What they have in common is turning the full history of an AI system into a traceable record. For document governance, this means you have to keep: where each category of data is processed and on what legal basis, which models (and versions) were called, and logs of where sensitive data flowed. This isn't for show; it's the evidence you can point to item by item when an auditor sits down to go through the books.

A pragmatic judgment: data sovereignty isn't something you can patch in by writing a compliance document right before launch; it's an architectural decision. Where embedding runs, where inference happens, and whose hands the chunks pass through are all essentially locked in the moment you choose a platform and draw your data flows. Thinking this through in advance as part of the ingestion process is far cheaper than reworking everything when the legal team comes knocking. The next section moves into a more dynamic scenario: when the knowledge base is no longer read-only and AI agents are allowed to write to it, how do audit and observability keep up?

When AI agents can write to the knowledge base: audit, tiering, and observability

The previous six steps rested on an implicit assumption: the knowledge base is read-only; people write into it, and machines read from it. But that assumption is collapsing. Many collaboration platforms now allow AI agents to write content directly into the knowledge base, modify entries, and generate new documents, turning the knowledge base from a passively searched corpus into a system that grows on its own. Once the write side is opened up, all the earlier effort on cleaning, structuring, and freshness can be quietly diluted. An AI agent batch-fills three hundred documents in the middle of the night, nobody knows which version of the data, which model, or which prompt it relied on, and the next second that content becomes the "facts" someone else retrieves.

So this section isn't about how to feed content in but about what your backstop is when machines can also put things in. The core is three complementary pieces: operation audit that can trace every write, write permissions tiered by role, and a clearly defined boundary of autonomy, spelling out which operations an AI agent can complete on its own and which must pause and wait for human confirmation. Without any one of these, quality will degrade in a way that throws no errors and raises no alerts.

We're already tuning chunk size and swapping embedding models. Why are results still unstable?

Because you're tuning the retrieval layer, while the root of the instability is often in the data layer. Chunk size and choice of embedding model determine "whether you can find the right thing in a pile of content," but if that pile itself contains duplicate versions, expired entries, and passages silently rewritten by AI agents, no matter how precisely you chunk, what gets retrieved is still dirty. A typical signal: the same question is answered correctly today and incorrectly two weeks later, with nobody having touched the retrieval configuration in between. Most likely, the underlying documents were written to or drifted. I'd suggest doing one thing first: attach to every model output the data version and model version it actually cited, and make that a traceable audit record. That way, the next time results get worse, you can pinpoint directly which piece of data changed and when, instead of repeatedly trial-and-erroring between chunk sizes and embedding models. Retrieval tuning has an upper bound; data observability is what sets the ceiling.

Should permission filtering happen at ingestion or at query time?

At query time. Splitting by permissions at ingestion means maintaining a separate copy for every access level, and every time a document is updated you have to sync multiple copies, which will fall out of alignment sooner or later. The right approach is to store a single copy of the content with permission labels at ingestion and, at the moment of retrieval, filter the candidate chunks using the actual permissions of the user running the query. Whatever the user can't see never enters the context sent to the model. This is especially critical once the knowledge base is writable: new content written by AI agents must also carry permission metadata and go through the same filter; otherwise, a chunk that should have been restricted, once generated by an agent, can leak out around the existing access controls. Permissions aren't a one-time action at ingestion; they're a real-time judgment at query time.

What specifically does pre-ingestion cleaning involve?

In order of priority, it roughly breaks down into these layers. The first layer is removing dirt: eliminating duplicate documents, merging contradictory versions, and stripping out formatting noise (headers and footers, leftover navigation, garbled characters), so that each fact has a single authoritative statement. The second layer is structuring: extracting unstructured content like PDFs, chat logs, and tickets into governable data with fields, with title, ownership, time, and owner all captured as metadata, so that later freshness governance and permission filtering have something to grip. The third layer is security scanning: running sensitive information and content security checks before content enters the database and keeping reviewable examples of the detection results, rather than discovering problems only when they're retrieved. Once you reach the stage where AI agents can write, all three layers have to be applied again on the write side. Human entry and machine generation are held to the same admission standard; nothing is exempt from inspection just because an agent wrote it.

How do you align cross-border data compliance with a RAG deployment?

The key is to be clear that what gets sent out isn't the entire document but the retrieved chunks, so compliance judgments need to be made at the chunk level. Before a chunk leaves your perimeter and enters an external model, confirm at least three things: whether the network layer is isolated, whether the content being sent has gone through redaction and content security controls, and whether a traceable audit log is kept that can show when this chunk was sent where, and because of whose query. One more thing is easy to miss: deletion must be verifiable. When a document must be taken down for compliance reasons, you need a tested deletion process that confirms it has been removed not only from the original storage but also completely purged from indexes, caches, and any derived content AI agents may have written, and you need to be able to produce evidence that the deletion took effect.

Put all of this together and governance is no longer a one-time ingestion action but a closed loop: use observability tools to continuously track how the model performs in real-world scenarios, and record every prompt involved, model version, test result, issue found, and corresponding improvement. In the next round of optimization, what you change is a specific, documented step, not parameters retuned by gut feel. Once machines can write to the knowledge base, this record is the only reliable buffer between you and quality drift.