Knowledge MeshHow Knowledge Mesh works inside one spoken turn
Retrieval, resolution and governance, all of it finished before the agent begins to speak.
Voicing Team
9 min read
Accurate, current, well grounded, and a document this caller had no right to hear. The agent has already said it out loud.
Voicing Team8 min read
Knowledge MeshThe retrieval worked perfectly.
The caller asked a question. The system searched the knowledge base, found the most relevant document, ranked it first, and grounded the agent’s answer in it with a clean citation trail. If you were evaluating this on retrieval quality, precision, recall, grounding fidelity, you’d score well.
The document was a claims note on a policy held by the caller’s spouse.
And the agent has already read the relevant part aloud.
That’s the whole problem, compressed into one paragraph. Every part of the retrieval pipeline did its job. The failure was that nobody asked whether this particular caller was permitted to hear this particular thing, and by the time anyone could have asked, the words were already in the air.
Access control in retrieval-augmented systems is a known problem, and the text world has developed a workable set of answers. Filter at query time. Check permissions on results before rendering. Redact before display. Some systems even generate first and validate second, produce a response, evaluate it against policy, and suppress it if it fails.
That last pattern is where voice breaks the model entirely.
A text agent that surfaces something it shouldn’t can have the message withheld. A voice agent that says it has already disclosed it. The disclosure occurred at the moment of speech synthesis. There is no queue to intercept, no render step to block, no message to delete. The audio left the system and entered a human ear.
Which means the entire class of mitigations that operate after generation is unavailable to you. Output filtering is not a control in voice. It is, at best, a way of finding out that a breach happened.
Three more voice-specific aggravations, all of which sound minor and aren’t:
In voice there is no interval between generation and delivery. Speech synthesis is the disclosure.
Speech is a disclosure event with no draft state. In text there is a meaningful moment between “the model produced this” and “the user saw this.” In voice, those collapse. The synthesis is the delivery.
Callers are harder to authenticate than sessions. A text user is behind an authenticated session with a known identity. A caller is a voice on a phone line claiming to be someone, and identity verification is something your agent does during the conversation, which means there is a window at the start of every call where the agent is talking to an unverified party. Anything retrievable in that window is retrievable by someone who has not yet proven who they are.
Recordings make it durable. The disclosure isn’t only the moment. It’s in the call recording, the transcript, and the post-call analysis, propagating into systems with their own retention policies. One unpermitted retrieval becomes an artifact in four places.
Draw the standard pipeline and the permission check tends to land in one of three places.
After generation. The agent produces a response, a policy layer inspects it, and unacceptable output is blocked. Works in text. In voice, as established, this is detection rather than prevention.
After retrieval, before the model. Better. Documents are fetched, then filtered against the caller’s permissions, then what survives goes into context. This is where a lot of enterprise RAG lands and it’s a genuine improvement, but it has a subtle failure mode. The retrieval ranking has already been computed over documents the caller’s system says the agent is left grounding on the fourth-best document while behaving as though it has the best answer. Confidence stays high; quality quietly drops.
Inside retrieval itself. The permission context is part of the query itself. Documents the caller isn’t entitled to see are not candidates. They aren’t retrieved, ranked, filtered, or held in a buffer. They’re outside the search space for this conversation.
Only the third option makes the guarantee you actually need, which is not “the agent didn’t say it” but “the agent never had it.”
That’s a meaningfully stronger claim, and it’s the one a compliance function will care about, because it’s the difference between a control that depends on the model behaving correctly and a control that doesn’t depend on the model at all.
For permission-aware retrieval to work, the retrieval layer needs to know things at the moment of the query that most knowledge bases have no concept of.
Who is on the call, and how sure are we? Not just an asserted identity, a verification state. Pre-verification and post-verification are different permission contexts, and the agent moves between them mid-conversation. The retrieval layer has to move with it.
What is this caller entitled to? Their own records, obviously. But entitlement in enterprise relationships is layered: an authorised representative on a corporate account, a power of attorney, a parent on a minor’s policy, a broker with delegated authority over some fields and not others. This is data you almost certainly already hold, and almost certainly haven’t connected to your voice agent’s knowledge retrieval.
What is the sensitivity of the content itself? Independent of who’s asking. Some material is restricted regardless of entitlement, internal notes, fraud flags, litigation holds, clinical detail requiring a different disclosure pathway.
What jurisdiction and what purpose? The same retrieval can be permitted for one purpose and not another, and permitted in one jurisdiction and not another. A regulated multinational contact centre lives with this constantly.
Has policy changed since this document was indexed? Permissions are not static. A control that was evaluated at indexing time is a control that reflects last quarter’s policy.
That last point is the argument for runtime rather than build-time enforcement. Permissions applied when content was ingested are a snapshot. Permissions applied at the moment of retrieval reflect current state, the current verification status of this caller, the current entitlement record, the current policy.
Knowledge Mesh applies policy, permission and contextual rules to retrieved context before it reaches the agent, blocking, filtering or constraining content based on sensitivity, policy and domain requirements at the moment of retrieval rather than at index time or after generation.
Two deliberate limits on that claim, because this is a domain where overclaiming is how vendors end up in difficult conversations.
This is a retrieval control, not a complete data governance programme. It governs what the agent can reach and cite during a conversation. It does not replace access control in your systems of record, your identity infrastructure, or your data classification work. It consumes those things. If your entitlement data is wrong, permission-aware retrieval faithfully enforces the wrong entitlements.
Enforcement quality depends on the fidelity of what it’s given. A retrieval layer can only enforce distinctions that exist in the data it receives. If your systems don’t distinguish between a policyholder and an authorised representative, neither can we. Getting this right in a regulated deployment is partly an integration exercise on your side, and any vendor who tells you it’s purely their software is setting you up for disappointment.
In a technical evaluation, don’t ask whether the platform supports access controls. Everyone says yes, and everyone is telling a version of the truth.
Ask this instead:
At what point in the pipeline is the permission decision made, and can you show me that unpermitted content was never retrieved, rather than that it was retrieved and then filtered?
The distinction is auditable. Retrieval logs will show you which documents entered the candidate set. If restricted material appears there and was subsequently filtered, the control is a filter and it depends on the filter working. If restricted material never appears, the control is structural.
Then ask the voice-specific follow-up, because this is the one that gets skipped:
What can the agent retrieve before the caller has been identity-verified?
Every call has that window. Most teams have never specified what is reachable inside it. It is, in our experience, the most common gap in otherwise well-governed voice deployments, and it’s the one that’s hardest to explain afterwards, because the honest answer is that nobody thought about the first thirty seconds.

Voice infrastructure on the contact centre floor
A working session with an engineer who has deployed inside a bank’s perimeter. We map your telephony, data boundary and handoff rules, and tell you what we would not automate.