Knowledge MeshHow Knowledge Mesh works inside one spoken turn
Retrieval, resolution and governance, all of it finished before the agent begins to speak.
Voicing Team
9 min read
Everything written about context rot assumes an agent that can pause, compact and retry. A caller is breathing on the other end of the line.
Voicing Team8 min read
Knowledge MeshAt turn four, the agent is excellent.
It’s crisp, it’s on task, it remembers the account number the caller gave at the start, and it handles a mid-sentence correction without stumbling. If you were demoing this call to a buyer, you’d stop here.
At turn twenty-two, something has gone wrong. The agent asks a question the caller already answered. Then it makes a statement that contradicts something established eleven turns ago. Then, and this is the one that generates the complaint, it confidently repeats a detail the caller explicitly corrected.
Nothing changed. No deploy, no config edit, no model swap. The context window didn’t overflow; you’re nowhere near the limit. The agent simply got worse as the conversation got longer.
This is context rot, and by now there’s a substantial literature on it. There are benchmark studies, arXiv papers, engineering blog posts from frontier labs, and a growing toolkit of mitigations. Almost none of it applies to you.
Read the published work on context rot and you’ll notice something once you’re looking for it. Nearly every example is a coding agent.
An agent working through a repository. An agent doing deep research across forty documents. An agent on a long-horizon software task, twenty messages into a session, starting to contradict decisions it made an hour earlier.
The findings themselves are solid and they generalise fine. As context grows, models attend to it less reliably, the effective window, where a model actually performs well, is meaningfully smaller than the advertised token limit. Information that is technically present in the window still gets missed, and distractor content makes it worse. Context is a finite resource with diminishing marginal returns, not a bucket you fill.
Dead air on a phone call is not a pause for thought. It reads as a malfunction, and the caller reacts to it.
It’s the mitigations that don’t generalise. The standard toolkit is compaction, structured note-taking, and sub-agent isolation. Summarise the conversation so far and reinitialise with the summary. Have the agent write notes to an external file. Delegate a messy subtask to a sub-agent and bring back only the clean result.
Every one of those requires something a voice agent does not have: permission to stop and think.
A coding agent can spend four seconds compacting its own history. Nobody notices. There is no one waiting; the developer is in another tab.
A voice agent has a live human being on the other end of a phone line, and the constraints that follow from that are not incremental.
There is no pause. Dead air in a phone conversation is not neutral. It reads as a malfunction. A caller experiences two seconds of silence as the system breaking. You cannot run a summarisation pass between turns and hope nobody minds.
There is no retry. A text agent that produces a bad answer can be regenerated before anyone sees it. Some systems do this routinely, generate, evaluate, discard, regenerate. A voice agent’s output is synthesised into speech and transmitted. Once it has been spoken, it has been said. There is no discard stage.
There is no scrollback. A caller cannot re-read what the agent said six turns ago to resolve a contradiction. They only have their own memory of the conversation, which means every inconsistency lands as the system is wrong rather than let me check what it said before.
The turn budget is unforgiving. Whatever context assembly happens has to happen inside the gap between the caller finishing their sentence and the agent beginning its reply. That budget is short enough that any strategy involving a second model to clean up context is competing directly with the thing the caller actually notices.
Put those together and you get the core asymmetry:
In an async agent, context rot is a quality problem you can clean up after it appears. In a voice agent, context rot is a live failure you have to prevent before the turn begins.
The general failure modes are documented. Here’s what each one actually looks like on a phone call, which is where they stop being abstract.
Poisoning. A wrong fact enters context and gets referenced repeatedly. In a coding agent this produces a bug you eventually notice under legal review. On a call, the agent misheard a digit in an account number in turn two, and has now confidently cited that wrong account four times, and the caller is starting to wonder whether they’re talking to a system that has their records mixed up with somebody else’s. Recovery is harder because the wrong value is now load-bearing for everything after it.
Confusion. Too much irrelevant material, often too many overlapping tool definitions, and the agent starts choosing badly. Models degrade measurably as tool count rises, and overlapping tool descriptions make it worse even when every tool is relevant. On a call this presents as the agent doing something almost-right: querying the balance tool when it needed the payment-history tool, and then answering a question the caller didn’t ask.
Rot. The general decline across a long session. This is the turn-four-to-turn-twenty-two story at the top of this post. It’s the most common one in voice because collections calls, claims calls and support escalations are long, and length is the variable that drives it.
Clash. Conflicting information inside the same context, usually because an early incorrect attempt is still sitting in the history alongside its correction. The agent has both the wrong value and the right one and no strong basis for preferring either. On a call this is the agent contradicting itself within thirty seconds, which is the single fastest route to a caller asking for a human.
Notice that all four are worse in voice for the same underlying reason: the caller is a participant in the failure, in real time, and their reaction becomes part of the context too.
If the standard mitigations assume async, what’s left?
The answer is that the work has to move earlier, out of the turn and into the retrieval layer. You cannot clean up context mid-call, so the discipline becomes never letting the window get into a state that requires cleaning.
Retrieve less, more precisely. The instinct with retrieval is to be generous: pull twelve chunks, let the model sort it out. That instinct is actively harmful here, because every irrelevant chunk is both a distractor and a latency charge. Just-in-time retrieval of a small number of high-signal items beats pre-loading a large context and hoping the model attends to the right part of it.
Resolve state outside the window. A conversation has facts that are settled, identity verified, intent classified, account confirmed, correction applied. Those don’t need to live in the transcript history where they compete for attention with everything else. Held as resolved state and injected as a compact, current summary of what is true, they stop being subject to rot.
Make correction destructive. When a caller corrects something, the old value should not remain in context alongside the new one. This is the direct fix for clash, and it’s a design decision rather than a model capability, the retrieval layer has to be able to replace rather than append.
Isolate the layers when diagnosing. A significant share of apparent context problems are audio problems, and vice versa. Being able to run the same conversation logic with speech recognition out of the path is what turns “is this a prompt problem or a transcription problem?” from a debate into a test. We wrote about why that matters for deployment speed in A Draft Takes a Minute, Production Takes a Month.
Everything in that list has one property in common: it happens before the model sees anything.
That’s why we don’t think of context management as a feature of our voice agents. It’s a layer they’re built against. Knowledge Mesh decides what enters the window on each turn, retrieving narrowly, holding resolved state outside the transcript, and replacing rather than accumulating when facts change.
The alternative architecture, an agent that exists first, with retrieval wired in afterwards, inherits the async playbook by default, because it was designed for a world where you could clean up later. Then it gets deployed into voice, where you can’t, and the failure shows up somewhere around turn twenty-two on a real customer call.
We’ve written about why the agents were engineered against the layer rather than the other way around in Ten Specialists, Not One Assistant, and about how the layer works in How Knowledge Mesh Actually Works.
You don’t need to buy anything to find out whether you have this problem.
Take your last thousand calls. Bucket them by turn count. Then measure your containment, or your resolution rate, or your escalation rate, whichever you trust most, within each bucket.
If the metric is flat across turn counts, your context handling is fine and you can stop reading.
If it degrades as turn count rises, you have context rot. And the important part: the degradation curve tells you roughly where your effective limit sits, which is almost certainly much earlier than anyone assumed when the agent was configured.
Most teams have never plotted this, because the aggregate number looks fine. The aggregate averages your short calls, which work beautifully, with your long calls, which are where your hardest and most valuable conversations live.

Voice infrastructure on the contact centre floor
A working session with an engineer who has deployed inside a bank’s perimeter. We map your telephony, data boundary and handoff rules, and tell you what we would not automate.