Skip to content

Blog

Retrieval Is a Permissions Problem First

Accurate, current, well grounded, and a document this caller had no right to hear. The agent has already said it out loud.


Voicing Team8 min read

Contents
A filing cabinet with one drawer open, one drawer padlocked and a rope barrier in front, a handset resting on top, drawn in fine ink linesKnowledge Mesh
The document01 / 02

The retrieval worked perfectly.

The caller asked a question. The system searched the knowledge base, found the most relevant document, ranked it first, and grounded the agent’s answer in it with a clean citation trail. If you were evaluating this on retrieval quality, precision, recall, grounding fidelity, you’d score well.

The document was a claims note on a policy held by the caller’s spouse.

And the agent has already read the relevant part aloud.

That’s the whole problem, compressed into one paragraph. Every part of the retrieval pipeline did its job. The failure was that nobody asked whether this particular caller was permitted to hear this particular thing, and by the time anyone could have asked, the words were already in the air.

Why voice makes this a different category of failure

Access control in retrieval-augmented systems is a known problem, and the text world has developed a workable set of answers. Filter at query time. Check permissions on results before rendering. Redact before display. Some systems even generate first and validate second, produce a response, evaluate it against policy, and suppress it if it fails.

That last pattern is where voice breaks the model entirely.

A text agent that surfaces something it shouldn’t can have the message withheld. A voice agent that says it has already disclosed it. The disclosure occurred at the moment of speech synthesis. There is no queue to intercept, no render step to block, no message to delete. The audio left the system and entered a human ear.

Which means the entire class of mitigations that operate after generation is unavailable to you. Output filtering is not a control in voice. It is, at best, a way of finding out that a breach happened.

Three more voice-specific aggravations, all of which sound minor and aren’t:

In voice there is no interval between generation and delivery. Speech synthesis is the disclosure.

Speech is a disclosure event with no draft state. In text there is a meaningful moment between “the model produced this” and “the user saw this.” In voice, those collapse. The synthesis is the delivery.

Callers are harder to authenticate than sessions. A text user is behind an authenticated session with a known identity. A caller is a voice on a phone line claiming to be someone, and identity verification is something your agent does during the conversation, which means there is a window at the start of every call where the agent is talking to an unverified party. Anything retrievable in that window is retrievable by someone who has not yet proven who they are.

Recordings make it durable. The disclosure isn’t only the moment. It’s in the call recording, the transcript, and the post-call analysis, propagating into systems with their own retention policies. One unpermitted retrieval becomes an artifact in four places.

Where most architectures put the check, and why that’s wrong

Draw the standard pipeline and the permission check tends to land in one of three places.

After generation. The agent produces a response, a policy layer inspects it, and unacceptable output is blocked. Works in text. In voice, as established, this is detection rather than prevention.

After retrieval, before the model. Better. Documents are fetched, then filtered against the caller’s permissions, then what survives goes into context. This is where a lot of enterprise RAG lands and it’s a genuine improvement, but it has a subtle failure mode. The retrieval ranking has already been computed over documents the caller’s system says the agent is left grounding on the fourth-best document while behaving as though it has the best answer. Confidence stays high; quality quietly drops.

Inside retrieval itself. The permission context is part of the query itself. Documents the caller isn’t entitled to see are not candidates. They aren’t retrieved, ranked, filtered, or held in a buffer. They’re outside the search space for this conversation.

Only the third option makes the guarantee you actually need, which is not “the agent didn’t say it” but “the agent never had it.”

That’s a meaningfully stronger claim, and it’s the one a compliance function will care about, because it’s the difference between a control that depends on the model behaving correctly and a control that doesn’t depend on the model at all.

Runtime guardrails, and what they have to know

For permission-aware retrieval to work, the retrieval layer needs to know things at the moment of the query that most knowledge bases have no concept of.

Who is on the call, and how sure are we? Not just an asserted identity, a verification state. Pre-verification and post-verification are different permission contexts, and the agent moves between them mid-conversation. The retrieval layer has to move with it.

What is this caller entitled to? Their own records, obviously. But entitlement in enterprise relationships is layered: an authorised representative on a corporate account, a power of attorney, a parent on a minor’s policy, a broker with delegated authority over some fields and not others. This is data you almost certainly already hold, and almost certainly haven’t connected to your voice agent’s knowledge retrieval.

What is the sensitivity of the content itself? Independent of who’s asking. Some material is restricted regardless of entitlement, internal notes, fraud flags, litigation holds, clinical detail requiring a different disclosure pathway.

What jurisdiction and what purpose? The same retrieval can be permitted for one purpose and not another, and permitted in one jurisdiction and not another. A regulated multinational contact centre lives with this constantly.

Has policy changed since this document was indexed? Permissions are not static. A control that was evaluated at indexing time is a control that reflects last quarter’s policy.

That last point is the argument for runtime rather than build-time enforcement. Permissions applied when content was ingested are a snapshot. Permissions applied at the moment of retrieval reflect current state, the current verification status of this caller, the current entitlement record, the current policy.

What we built, stated narrowly

Knowledge Mesh applies policy, permission and contextual rules to retrieved context before it reaches the agent, blocking, filtering or constraining content based on sensitivity, policy and domain requirements at the moment of retrieval rather than at index time or after generation.

Two deliberate limits on that claim, because this is a domain where overclaiming is how vendors end up in difficult conversations.

This is a retrieval control, not a complete data governance programme. It governs what the agent can reach and cite during a conversation. It does not replace access control in your systems of record, your identity infrastructure, or your data classification work. It consumes those things. If your entitlement data is wrong, permission-aware retrieval faithfully enforces the wrong entitlements.

Enforcement quality depends on the fidelity of what it’s given. A retrieval layer can only enforce distinctions that exist in the data it receives. If your systems don’t distinguish between a policyholder and an authorised representative, neither can we. Getting this right in a regulated deployment is partly an integration exercise on your side, and any vendor who tells you it’s purely their software is setting you up for disappointment.

The question that actually separates vendors

In a technical evaluation, don’t ask whether the platform supports access controls. Everyone says yes, and everyone is telling a version of the truth.

Ask this instead:

At what point in the pipeline is the permission decision made, and can you show me that unpermitted content was never retrieved, rather than that it was retrieved and then filtered?

The distinction is auditable. Retrieval logs will show you which documents entered the candidate set. If restricted material appears there and was subsequently filtered, the control is a filter and it depends on the filter working. If restricted material never appears, the control is structural.

Then ask the voice-specific follow-up, because this is the one that gets skipped:

What can the agent retrieve before the caller has been identity-verified?

Every call has that window. Most teams have never specified what is reachable inside it. It is, in our experience, the most common gap in otherwise well-governed voice deployments, and it’s the one that’s hardest to explain afterwards, because the honest answer is that nobody thought about the first thirty seconds.

What is permission-aware retrieval?
Retrieval where the requester’s permissions form part of the query rather than a filter applied to results. Content the requester isn’t entitled to access is never a retrieval candidate, so it is not fetched, ranked, or placed in the model’s context at any point.
Why isn’t output filtering sufficient for voice AI compliance?
Because output filtering intercepts a response before delivery, and in voice there is no interval between generation and delivery, speech synthesis is the disclosure. Once audio has been transmitted, it cannot be withheld. Post-generation controls in voice detect breaches rather than prevent them.
What’s the risk window at the start of a voice call?
Identity verification happens during the conversation, so every call begins with the agent talking to an unverified party. Whatever the retrieval layer can reach during that period is reachable by someone who has not yet proven who they are. Specifying pre-verification retrieval scope explicitly is a control most deployments overlook.
Does permission-aware retrieval replace our existing access controls?
No. It governs what a conversational agent can reach and cite at runtime, and it consumes entitlement and classification data from your existing systems. It enforces the distinctions present in the data it’s given, so it complements systems-of-record access control rather than substituting for it.

Read nextArticle

Translation belongs in the media path, not beside it

9 min readRead the article

Closing02 / 02

Bring one call type. Leave with an architecture.

An airport service desk at sunrise: a traveller with a suitcase at the counter, and an agent in a headset answering behind it.
Voicing

Voice infrastructure on the contact centre floor

A working session with an engineer who has deployed inside a bank’s perimeter. We map your telephony, data boundary and handoff rules, and tell you what we would not automate.