2026-06-19
Access control for RAG knowledge bases: keep AI from answering with documents users shouldn't see
In a RAG system, an unauthorized answer isn't a "wrong answer" — it means documents that should never have been exposed were fed to the model. Starting from the risk itself, this article walks through how to implement RAG permission isolation correctly: why filtering must happen at the retrieval layer, the correct pipeline order of retrieval → ACL filtering → Rerank, how to choose between physical and logical isolation, rules for passing tenant_id securely in multi-tenant systems, and how to close side channels and citation leaks — ending with a practical audit and acceptance plan.
An unauthorized answer isn't a wrong answer — it means restricted content was fed into the context
First, let's define the problem clearly. When a RAG knowledge base goes wrong, it usually falls into one of two categories. In the first, the retrieved content is limited and the model's answer isn't accurate enough — a quality problem you can grind down over time with retrieval optimization, prompt engineering, and human evaluation. The second is entirely different: the model answers very accurately, so accurately that it recites, word for word, a document you were never authorized to see. That is not a quality problem; it is a security incident. Lumping the two together is the first pitfall many teams only recognize after going live.
Consider a few concrete scenarios, none of them rare in practice. An ordinary employee types "What strategic changes did management decide on this year?" into the internal assistant; the system retrieves minutes from a closed-door executive meeting from the vector database, and the model dutifully summarizes them. A colleague outside finance asks "Roughly what's our team's budget for next year?" and gets the neighboring department's budget in the same answer. An outsourced on-site contractor with only a regular account manages to extract the design documents for a core R&D module. In a multi-tenant SaaS product, an operations user at customer A asks a question in the back office and hits a contract template uploaded by customer B. What these scenarios have in common: the person asking made no "attack" at all. They simply typed as usual, and the system itself handed over content it should never have provided.
There is a key judgment here about the order of events that deserves to be singled out. Once unauthorized content enters the model's context window, the leak has already happened in the engineering sense — regardless of whether the model ultimately outputs it. Many people instinctively try to remedy this on the output side: adding a sensitive-word filter, having the model "declare" before answering that it can only discuss what the user is authorized to see, or having another model review the output. All of these come too late. As soon as that text is assembled into the prompt, it participates in that round of reasoning: the model might paraphrase it in its answer, might give it up after a few follow-up questions, might expose the document title in the citation list, or might even reproduce it in later turns because of residual context in a multi-turn conversation. What you filter out is only one phrasing in one output; you cannot undo the fact that the information itself is already present.
So there is only one real control point: before the model generates. Specifically, permissions must be resolved during the two steps of retrieval and context assembly, so that unauthorized documents never make it into the candidate set, let alone the prompt. This is structural security, not filtering that relies on luck. Putting permissions after generation is like opening the safe, letting someone take a look, and then debating whether to let them describe what's inside — once the order is wrong, everything that follows is damage control.
Then can you write permission rules into the system prompt and let the model decide for itself "whether you're allowed to see this passage"? No, and the reason lies in the same timeline: for the model to make that judgment, the content must already be in its context, so the act of judging itself happens after the leak. And even setting order aside, asking a model to enforce access control means handing a deterministic logical decision to a probabilistic system. It may comply today and be talked around by a cleverly worded follow-up tomorrow. Permission checks require a definitive "either grant or deny" result; that job belongs to code and the data layer, not to the model's self-discipline.
With this understanding in place, the engineering design that follows has a foundation. In RAG, access control is not an optional enhancement but a first-class citizen of the retrieval pipeline: it determines which documents are eligible to become candidates, whether they can enter ranking, and whether the model ultimately gets to see them. From this section onward, every approach we discuss — the ordering of retrieval-layer filtering, the trade-offs between physical and logical isolation, the source of identity in multi-tenant systems, caching strategy, closing side channels — is essentially answering the same question: how do you guarantee that, before the model starts writing, every piece of material laid out in front of it is something the person asking is genuinely authorized to see? This criterion runs through the entire article; whenever you encounter an approach, measure it against this test first.
Pin permissions to the retrieval layer: why "filter after you get the results" isn't secure
Let's start with an equation that is easy to get wrong in engineering: many teams assume that "the system has login and distinguishes admins from regular users, so permissions are done." This equates two different layers. Login authentication answers "who you are and whether you can enter this feature" — it is an access gate. Knowledge base permissions must answer "can this specific document chunk influence the answer for this particular person?" The former operates at feature granularity, the latter at data granularity, and the entire retrieval semantics lies between them. An open door does not mean every document behind it is open to you.
The real divergence appears in the timing of filtering. One approach is to retrieve normally, take the matching chunks, and then filter them by permission before returning — it sounds like the same result, since only what the user is authorized to see is exposed externally. But this path has cracks in engineering terms. The problem lies not in the final display but in the intermediate step: the filtered-out chunks have already been pulled from the vector database, and their existence, their relevance scores, even their titles, have already entered the system's processing chain. As soon as any subsequent step (reranking, logging, debug output, cache keys) touches this unfiltered batch of candidates, unauthorized content has a way out. "Filtering after the fact" protects the exit, not the process.
So filtering must move ahead of the two actions of retrieval and context injection. It becomes clearer from another angle: what the model can ultimately say depends entirely on what was fed into its context. The only reliable way to ensure that unauthorized documents can never appear in an answer is to give them no chance to enter the context at all — that is, to make permission constraints part of the query when constructing retrieval conditions, so that unauthorized chunks are not even eligible to be fetched. This moves the protection boundary from the "output layer" to the "data acquisition layer"; the earlier it sits, the fewer steps there are that can be bypassed.
Structurally, the more robust approach is to extract permissions into an independent layer that sits above the business logic, rather than scattering it across individual call sites. Whether it's a RAG retrieval request or a tool call or knowledge query initiated from an AI agent conversation, every entry point passes this layer's check first and receives a visibility scope already narrowed by permissions before proceeding to retrieval. Making permissions a shared checkpoint up front, rather than something each feature implements separately, has two benefits: first, there is only one copy of the decision logic, so you won't end up with inconsistencies like "the retrieval API checks permissions but the agent path forgot"; second, auditing has a single point of record: who fetched which chunks under what permissions can be recorded and replayed centrally.
It should be stressed that the application layer's "admin/regular user" role separation is not redundant — it is a necessary first gate, responsible for the coarse-grained division of feature visibility and operation permissions. But it handles "can you use this feature," not "which of the dozens of chunks retrieved this time should you be allowed to see." The latter is a decision made down to the individual data item, dynamically and in combination with query semantics, and it must be owned by the retrieval layer itself. Only by keeping these two layers distinct can you avoid problems like "the roles were configured correctly, yet the knowledge base still answered with another department's contract terms."
In practice, when evaluating whether a RAG system's permissions can be trusted, don't just check whether it has login and roles; look at which step of the pipeline the filtering happens in. If the answer is "after retrieval, before return," it protects the display, not the data; only if the answer is "permission constraints are built into the retrieval conditions and unauthorized chunks are never fetched" have permissions been pinned to the retrieval layer. This line is the boundary between "looks secure" and "is actually secure."
The correct order of the retrieval pipeline: retrieval → ACL filtering → Rerank → top 5
The retrieval chain involves four actions: vector retrieval, permission filtering, reranking, and truncation to the top few results. Their relative order is not a matter of style; it directly determines whether unauthorized content leaks. Getting the order right matters more than tuning any model parameter. A workable default configuration: widen the retrieval stage to several dozen candidates (50 is a safe starting point), pass them to the permission service for precise item-by-item checks, filter out those the user cannot access, rerank what remains, and finally take the top five into the context. Below, we break down why each step sits where it does.
Cast a wide enough net at retrieval
Many people intuitively set the initial retrieval top_k to 5, on the grounds that "we only use five in the end anyway." That intuition breaks down once permission filtering is added. Retrieval ranks by vector similarity, which has nothing to do with whether the current user has permission — the most similar results may well be exactly the documents this user cannot see. If you fetch only 5 initially, you may be left with 1 or even 0 valid results after filtering, and the model answers with incomplete context: either vaguely, or by simply saying it doesn't know. The user's impression is "this useless system can't find anything," but the root cause isn't retrieval quality — the retrieval net was cast too narrowly, and permission filtering emptied it in one cut.
So the retrieval stage should be sized backward from "enough left over after filtering." If a department can see only a small share of the documents, the retrieval count needs to be raised accordingly. 50 is an empirical value, not an iron rule — you can monitor the filter pass rate and adjust dynamically: retrieve more for tenants with low pass rates, and trim a bit for tenants with high pass rates to save compute. The key is not to let filtering drain the candidate pool.
Filtering must sit before reranking
Placing permission filtering after reranking is an ordering mistake that looks harmless but is actually dangerous. The essence of reranking (Rerank) is having a model read the body text of candidate documents and rescore them by relevance to the question. Note the "read the body text" part — as long as you rerank first and filter second, the content of unauthorized documents has already been sent to the reranking model. Even if they are filtered out at the end and never enter the final context, the leak already happened the moment reranking ran.
If you use a self-hosted reranking service running on your internal network, this risk is still manageable — at least the data hasn't left your perimeter. But many teams call an external Rerank API for convenience, and that is a different matter entirely: the body text of unauthorized documents gets packaged and sent to a third-party provider's servers. This is no longer just logical unauthorized access; it is sending tenants' sensitive data outside your own boundary. In a compliance audit, this is the kind of issue that becomes a major incident — an entirely different order of magnitude from "the answer said one sentence too many."
Conversely, filtering before reranking keeps the logic clean: every candidate that enters the reranking model is one the current user is confirmed to be authorized to see. Reranking only orders legitimate candidates, and unauthorized content never gets a chance to be touched by any model, internal or external. It also saves compute along the way: reranking is billed and timed per candidate, so cutting out the unauthorized ones first means reranking only processes what will actually be used.
Treat this chain as a line of defense
Put together, these four steps form a funnel that narrows in one direction: 50 candidates → however many legitimate ones remain after permission filtering → reranking → top five. Each step only shrinks or reorders the data and never puts back anything already filtered out. This monotonically narrowing property is the engineering basis on which you can promise externally that "unauthorized content will not enter the context" — as long as the filtering decision is reliable and it sits before reranking, no unauthorized document can appear out of nowhere, however the reranker orders things or however many results are kept after truncation.
In implementation terms, we recommend routing the permission filtering step through a centralized permission service that checks every retrieved document individually, rather than stuffing a pile of scattered if-checks into the retrieval layer. That way the filtering logic lives in one place, is testable and auditable, and if the permission rules change one day, only one place needs updating. Among the check's inputs, the user identity and tenant identifier must come from trusted server-side sources — a point that must be nailed down especially in multi-tenant scenarios, but that is the subject of a later section.
Choosing among three isolation approaches: from dedicated indexes to row-level security
Permission isolation is not a switch but a continuous spectrum of strength. From an implementation standpoint, practical approaches fall roughly into three tiers: a dedicated index or database per tenant; shared storage with mandatory filtering on tenant_id and acl_tags; and row-level control that pushes permissions down into the data source itself. Strength runs from high to low, while cost and flexibility run the opposite way. Which tier you choose depends on your data sensitivity, the order of magnitude of your tenant count, and how much operational cost your team is willing to pay for isolation.
The strongest tier is physical isolation: each tenant gets its own vector index, its own database, even its own namespace. Its advantage is a boundary so clear it barely needs explaining — queries cannot touch anyone else's data, because that data simply isn't in the index the current connection points to. When something goes wrong, the scope of investigation is naturally confined to a single tenant. The cost is just as plain: once the tenant count scales up, indexes fragment badly, resource utilization drops, hot and long-tail tenants get mixed into the same scheduling, and ops has to handle versioning, backups, and migrations for thousands of small databases. This path suits scenarios with few tenants, large per-tenant data volumes, and hard compliance requirements — for example, a handful of large customers, each with a dedicated setup; it does not suit SaaS models with tens of thousands of small tenants.
The middle tier is logical isolation with strong constraints: all tenants' data lives in the same storage, every chunk carries labels such as tenant_id and acl_tags, and queries push these labels into the filter as mandatory conditions. It strikes a balance between resource efficiency and isolation and is the pragmatic choice for most multi-tenant systems. But this tier's security rests entirely on "the filter is never bypassed" — the moment a single query path forgets to include tenant_id, cross-tenant leakage occurs. So it places the highest demands on engineering discipline: filter conditions must be funneled through a single unified retrieval entry point, no business code may hand-write raw queries, and both writing and validating labels must be backed by tests. The hardware costs it saves you will be paid back in the form of code reviews and regression tests.
The tier closest to the data is row-level security, typically implemented with PostgreSQL's Row Level Security. It writes filtering rules into database policies, so no matter who issues a query or which business path it takes, the database automatically appends row-level conditions based on the current session identity. Its value lies in moving "don't forget to filter" out of the application layer — even if the application issues a query without a tenant condition, the database blocks unauthorized rows for you. For teams worried about business code growing and filter points becoming ever more scattered, RLS is a backstop guardrail. The prerequisite is that your retrieval stack actually goes through this database layer and that the session identity itself is set in a trustworthy way; if vector search bypasses the relational database and hits the index directly, RLS has no reach over that part.
These three tiers are not mutually exclusive. A common, robust combination is to carry the bulk with logical isolation, add a layer of RLS as a backstop, and carve out physical isolation separately for the handful of most sensitive tenants. The permission model on the relational database side needs to keep pace as well. A simple but effective practice is to put a role field in the user table and distinguish admin from regular user at the database level, so that "who can see the administrative scope and who can see only their own" holds in the data structure itself rather than in if-checks scattered everywhere. Going further, create a separate association table such as user_file that records each document uploaded to the knowledge base along with its owner and visibility scope. This table matters beyond permission checks: it lets admins maintain a unified document inventory and, when something goes wrong, trace back to "who uploaded this document and whom it was granted to." Only when permissions and the asset ledger are managed together do auditing and revocation have something to hold onto.
One temptation deserves a specific warning: don't try to replace your permission system with a knowledge graph. Graphs excel at expressing relationships between entities and seem naturally suited to describing "who can access what," so some people want to model access control as edges in the graph and determine permissions via graph queries. This path usually leads into a new kind of permissions hell: once relationships get complex, edge semantics bloat; inheritance, transitivity, and exception rules tangle together; and eventually no one can say why a particular user can see a particular document. Permission decisions need to be deterministic, enumerable, and auditable, whereas graph queries offer flexible but hard-to-exhaust reasoning. Mix the two in one layer and, when debugging, you lose both the clarity of the graph and the controllability of the permission system. Let the knowledge graph do what it's good at, namely semantic association, and leave the final ruling on permissions to an independent mechanism with explicit rules. Draw this boundary clearly, and the trouble you save later will far outweigh the few extra tables you wrote up front.
A practical selection rule: with few tenants and hard compliance requirements, prefer physical isolation; for multi-tenant SaaS, use logical isolation as the base plus RLS as the backstop; either way, ledgers such as role and user_file in the relational database are non-negotiable, as they are the foundation on which both permission checks and after-the-fact audits depend. The stronger the approach, the more expensive it is, but the cost of a single leak is usually far higher than these costs; leaving a little extra margin on the stronger side is rarely something you'll regret.
Strong multi-tenant isolation: tenant_id must come from the server, never from the front end
Multi-tenant incidents often happen not because the filtering logic was written incorrectly, but because the identity field the filter depends on can itself be forged. You can order retrieval, ACL, and Rerank meticulously, but as long as tenant_id is read from the request body, the whole pipeline is built on sand. An attacker doesn't need to bypass any filtering rule; they only need to change tenant_id in the JSON to someone else's value, and the filter will faithfully retrieve all of another tenant's documents for them. It's like handing the judge's pen to the defendant.
So this section is really about just one thing: the source of identity fields. Fields such as tenant_id, user_id, and department_id — the ones that determine "what you can see" — must be taken from the server-side authentication context, injected by the gateway or authentication middleware after the session has been validated; downstream services neither accept nor read fields of the same name carried in the request. In other words, the front end can tell you what it wants to query, but not who it is. "What I want to query" is a business parameter; "who I am" is an identity fact. Their trust levels are worlds apart, and passing them together in the same request body is a cardinal sin in engineering.
Why "the front end can pass it" is practically the same as "it will be forged"
Some will argue that the front end passes tenant_id only for convenience and that normal users won't tamper with it. That assumption doesn't hold even for internal systems, and for external-facing services it is a disaster. The request body is fully transparent to the client; changing a field costs no more than opening developer tools and pressing Enter a few times. Once tenant_id can be controlled by the front end, the attack surface degrades from "break authentication" to "guess a valid tenant ID," and tenant IDs are often auto-incrementing integers or enumerable short strings. This effectively downgrades strong isolation into a warning sign with no teeth. In multi-tenant scenarios, tenant_id should be bound to every retrieval request as a mandatory isolation field, and its value must come solely from the server-side authentication result — a point repeatedly emphasized in engineering practice, precisely because forging tenant_id or user_id leads directly to cross-tenant unauthorized reads.
In implementation terms, there are a few actionable rules. First, after the authentication middleware parses the token, it writes the tenant and user identity into a request-scoped context object; the retrieval service reads values only from this context, and the code offers no entry point for "reading tenant_id from the body," eliminating misuse at the source. Second, push tenant_id down to the vector database or metadata store as a hard query condition, rather than querying first and comparing at the application layer — in the former, the database guarantees the boundary for you; in the latter, you have to remember to filter, and the two are not in the same league for reliability. Third, for every document ingested, fix its tenant ownership as an immutable metadata field at write time, to avoid later inferring ownership from "soft conventions" such as paths or namespaces.
Login authentication is a prerequisite for isolation to work
The premise of tenant_id coming from the server is that a trusted server-side identity exists in the first place. If the system allows anonymous access, the authentication context contains no tenant information at all, and strong isolation is out of the question. So both the knowledge base management pages and the chat pages should require login: only legitimate sessions that have passed account authentication may enter the system, with no anonymous entry points. This isn't just about keeping unauthorized people out; it's about ensuring that every retrieval request is backed by a clear, traceable identity, so the filtering logic has something to go on. A request without a login session should be rejected before it enters the retrieval flow, rather than proceeding to the filtering step with an empty identity and being special-cased there — fail-closed over fail-open is a non-negotiable default here.
The permission management plane must be locked down as well. Operations such as adding users, deleting users, and adjusting ownership and permissions should be restricted to a dedicated admin role; regular accounts can neither see nor change them. Persist user-to-tenant mappings and permission change records in reliable storage — first so that identity relationships aren't lost after service restarts or scaling, and second so that every permission change is on record. This may sound like basic back-office administration, but it is precisely the foundation of tenant_id's trustworthiness: the value injected by the server can be trusted because it is backed by a strictly controlled set of identity data.
A few easily overlooked corners
Even with multi-tenant isolation in place, a few gaps deserve close attention. The first is background batch jobs and asynchronous tasks. These calls often have no user session, and developers are tempted to take the shortcut of running over all data with superuser privileges, bypassing tenant boundaries in the process; we recommend assigning system tasks an explicit tenant scope or an explicit service identity too, rather than letting them through by default. The second is identity propagation across service calls. When service A calls service B with a user's identity and an intermediate step drops the context, service B may run the retrieval with a default or empty tenant. This kind of "identity lost in transit" problem is common in systems with long call chains and requires an agreed contract between services for passing identity. The third is cache keys: in a multi-tenant setup, every cache must include tenant_id as part of its key; otherwise one tenant's query results may be served to another tenant, and isolation silently fails at the cache layer.
The core position of this section can be summed up in one sentence: identity is a fact, not a parameter. Give the server exclusive write authority over tenant_id, combine it with mandatory login and controlled permission management, and multi-tenant isolation gains a fulcrum that cannot be forged. All the filtering, ranking, and auditing rest on this fulcrum — if it gives way, however elaborate the structure built on top, it's a castle in the air.
Performance and caching: batch checks, TTLs, and never caching highly sensitive data
Inserting ACL filtering into the retrieval path raises an unavoidable practical problem: every additional permission check adds another stretch of synchronous waiting. Retrieval already has to absorb the latency of vector search; if the filtering step adds a few hundred milliseconds on top, the perceived responsiveness of the whole chain collapses. This section covers how to make permission checks both accurate and fast, and which steps people skip for the sake of speed that actually cannot be skipped.
The easiest trap to fall into is item-by-item checking. It's normal for the retrieval stage to return dozens of candidate chunks at once; if you send a separate permission query for each chunk, that's dozens of round trips. Permission services are usually deployed independently, and while a single call doesn't look expensive, multiply it by the number of candidates and add network jitter, and tail latency gets ugly. Worse, this calling pattern amounts to an amplification attack on the permission service itself — a single retrieval request fans out into dozens of downstream requests, and as soon as concurrency rises, the downstream service falls over first.
The right approach is to consolidate the checks into one. Once you have the candidate set, deduplicate by document first — the same document often matches multiple chunks, and what really needs a permission check is the document, not every fragment. After deduplication, send this batch of document IDs to the permission service in one go and let it return, in bulk, whether each document is visible to the current principal. Dozens of candidates often shrink to a dozen or so documents after deduplication, which a single batch query can cover, cutting round trips from dozens to one. The precondition is that the permission service offers a batch API; if it only has a single-item API, that is the engineering gap to fill first — not something to paper over with a concurrent loop.
If you still want to cut latency after batching, caching comes next. In retrieval scenarios, results for similar questions within a short time overlap heavily, and a given user's visibility into a given set of documents basically doesn't change within a few dozen seconds, so these decisions are worth caching. The cache key must include the principal's identity and the document identifier — never cache by document alone. Caching by document alone means serving A's permission decision to B, which is no longer a performance optimization but unauthorized access.
Setting the TTL is essentially a balance between freshness and hit rate. Keeping the expiration on the order of 60 or 300 seconds is a fairly safe range: long enough to cover repeated retrieval within one continuous conversation, yet short enough that permission changes take effect within an acceptable window. The longer the TTL, the higher the hit rate, but also the longer the "leak-through window" after permissions are revoked — a user has already been removed from a project group, but the cache still thinks they have access, and during this time they can still get the content out by asking. So a longer TTL is not always better; it is a risk parameter you should be able to explain to your security team.
The real dividing line is highly sensitive data. For high-sensitivity permissions covering things like classified documents and compliance-restricted content, my position is: don't cache. The reason is simple: caching means accepting a staleness window, and high-sensitivity scenarios have near-zero tolerance for "permission revoked but still allowed through." Here it's better to run a real-time check every time and absorb the latency than to leave an exposure gap of several minutes for the sake of speed. Ordinary documents can trade caching for performance; highly sensitive documents trade performance for certainty. This is tiered handling by data classification, not a one-size-fits-all rule.
When implementing, string these points together: on the check path, deduplicate first and then batch; bind cache keys to the principal; set short TTLs for ordinary permissions; and mark high-sensitivity permissions as non-cacheable and force real-time lookups. Also leave one hook: when the permission system changes (membership adjustments, a document's classification level being raised), it's best to proactively invalidate the relevant cache entries rather than waiting for the TTL to expire naturally. If you can do proactive invalidation, the TTL for ordinary documents can be relaxed somewhat, improving both overall latency and freshness. The mechanism isn't complicated, but it decides whether permission filtering is merely "usable" or "production-ready."
Closing side channels and citation leaks: even if the answer doesn't leak, titles and hit counts can
The previous sections addressed "keeping content out of the context." But content is not the only path for permission leaks. A system that filters body text cleanly can still leak information through the cracks — and these cracks usually go unwatched until someone figures out the pattern after launch.
The most easily overlooked is the citation itself. The list of sources attached to a RAG answer is often treated in engineering as a "safe by-product that has already been filtered" and rendered as is. The problem is that the title line alone can be sensitive information. Take a document titled "Layoff List and Severance Plan (2025Q1)": even if the model doesn't quote a single word of its body, merely displaying this title in the citation list already tells people who shouldn't know that "the company is planning layoffs and the plan has been written up." File names, document summaries, the name of the space a document belongs to, its last modifier — all of this metadata is subject to permission checks and cannot be let through just because it isn't "body text."
Citation links need even stricter handling than the display. Filtering at display time doesn't mean it's safe when clicked. A common pitfall: list rendering goes through the ACL, but the click-through fetches the original from object storage directly by document ID, with no re-check on that hop. As a result, an unauthorized user who can't see the citation entry can still obtain a leaked link, or guess the ID pattern, and bypass the list to pull the full text directly. The right approach is to treat opening a link as an independent resource access request and re-run a document-level permission check on the server — a checkpoint independent of the one at retrieval time, with neither vouching for the other. Using non-enumerable random strings for IDs is something you should do while you're at it.
Truly insidious are statistical side channels. Systems often casually expose some "harmless" numbers in the interface: retrieval hit counts, similarity score distributions, messages like "Found N relevant documents for you." An attacker doesn't need to see any content: by repeatedly probing different keywords with a low-privilege account and watching the hit count jump from 0 to 1, they can infer that a certain name, project, or contract exists in the repository. This is a classic existence leak — confirming that "such a thing exists" is itself intelligence. Differences in error messages are equally fatal: if "access denied" and "not found" are returned differently, you are effectively telling the other party in plain terms that "the thing exists; you just can't see it." The two cases must look identical externally: either both return empty, or both return the same generic message, so that the shape of the response never varies with permission state.
Handling side channels requires a different mindset from handling body text. For body text, "filtering it out is enough"; for side channels, "authorized and unauthorized requests must look the same at the observable level." So auxiliary information such as hit counts and scores should either not be displayed or be recomputed from the filtered result set — never taken from the raw pre-filter numbers. There's an easy-to-fall-into counterexample here: for performance, computing and caching the total hit count first and then applying permission filtering. That cached number is pre-filter, so it still leaks. Statistics must always be computed downstream of the permission boundary.
Finally, there's auditability. It isn't compliance window dressing; it is the only real handle for investigating leaks. Side-channel leaks are covert, delayed, and hard to reproduce, and by the time you notice an anomaly, a long time has usually passed. So every retrieval must log the basis for its decisions: where the requester's identity and permission context came from, which policies were hit during filtering, how many results were originally retrieved, how many remained after filtering, and which document IDs were ultimately cited. With these numbers in hand, any later complaint that "so-and-so saw something they shouldn't have" can be investigated by pulling up the scene at the time and re-running it, confirming whether the filtering logic was wrong, the cache served stale data, or the permission data itself was misconfigured. Without this chain of evidence, you can't even say whether anything leaked, let alone pinpoint where. Making policies and the filtering process replayable is like buying after-the-fact insurance for your entire access control system.
Deployment and acceptance: audit fields, delivery targets, and replayable filtering evidence
At this point, whether your permission design succeeds no longer depends on the architecture diagram you drew, but on one thing: when something goes wrong, can you pull up the filtering process from that moment exactly as it happened? Acceptance testing a RAG permission scheme is essentially acceptance testing its observability — every retrieval must leave a trace that can be reviewed, rather than only checking whether the final answer is correct. The following four questions are the ones most often pressed in delivery reviews.
I've already written "do not leak confidential content" in the prompt. Do I still need retrieval-layer permissions?
Yes, and the two operate at different levels. Setting aside the debate over whether the model will obey, look at one engineering metric: how do you prove to the acceptance party that it works? A constraint in the prompt produces no recordable artifact — no fields, no counts, no replayable decision records. After the fact, the only way to reproduce anything is to re-ask questions and hope for the best, and that won't pass an audit.
Retrieval-layer permissions, by contrast, naturally produce evidence. After a request completes, the logs should contain three decreasing numbers: how many items were retrieved, how many remained after filtering against the access control list, and how many were ultimately assembled into the context — for example, 50 retrieved, 12 after filtering, 5 into the context — along with how many documents were checked in this request and how many were allowed through. This string of numbers is what the acceptance party really wants. It turns "did permissions take effect" from a matter of belief into a matter of fact that can be verified item by item. A prompt constraint can serve as a last-line fallback instruction, but it cannot replace retrieval-layer controls that leave an audit trail.
Should the permission check go before or after Rerank?
Before, no exceptions. Let's work backward from acceptance this time: the delivery target is "unauthorized evidence entering the context: 0," and to prove it, you must be able to produce a post-filter count that is fixed before reranking takes place. If the check comes after Rerank, unauthorized documents have already entered the scoring step — even if they are not selected in the end, their content has already been read by the external reranking service, which already constitutes a leak, and your logs have no way of recording that "it should never have been here."
So the ordering is not just a performance consideration; it directly determines whether the after_acl_count field means anything. Only when filtering comes first does this number represent "the size of the authorized candidate set," with reranking and truncation both taking place within that set, closing the chain of evidence.
Calling the permission service item by item is too slow. Can we add caching?
Yes, but cache hits must not break the audit chain. The common practice is to merge access decisions for a batch of document IDs into a single batch check and give the results a fairly short expiration time, so you don't hit the permission service for every item. This optimization was covered thoroughly in the performance section; here we add just one acceptance red line: even on a cache hit, the request must still record the number of documents checked and the number that passed, and the filtering reasons must still be replayable. Otherwise the cache becomes an audit blind spot — you investigate an unauthorized access incident, find "cache hit" in the logs, and can't retrieve which version of the permission snapshot the decision was based on. A log like that is as good as no log at all.
In addition, decisions for high-sensitivity content are never cached; they are computed fresh every time. Caching buys throughput at the cost of decisions lagging behind permission changes. That is acceptable for ordinary documents, but for highly sensitive documents you can't gamble on that lag.
If the answer body doesn't leak content, do citation titles and hit counts still count as leaks?
Yes. That is exactly why final_count must be recorded separately and logs must be broken down by dimension. For a document the user is not authorized to access, even if its body is never paraphrased, as soon as its title appears in the citation list — or the hit count quietly goes from the usual 3 to 4 — the other party has inferred that "there's something here I can't see." This kind of side-channel leak never shows up in the answer text; it can only be caught by auditing separately the different classes of logs: admin operations, RAG retrieval, and AI agent conversation runs.
The acceptance method is straightforward: take a low-privilege account and deliberately ask about content you know exists but that it isn't authorized to see, then check whether the returned citation titles and hit counts show any fluctuation. The three count fields plus categorized logs give this kind of probing a baseline for comparison.
Condense the four points above into an acceptance checklist: permissions must be enforced at the retrieval layer, zero unauthorized evidence enters the context, and every filtering action comes with a replayable decision reason. As for rollout pace, we recommend completing the permissions release in weeks two through four of the MVP — first put the audit fields and this line of defense against unauthorized access in place, then layer features on top. Get the order backward, and every retrieval entry point you add afterward is another floor built on a foundation with no audit trail.