In the Call, Not Beside It
Bolt-on translation sits parallel to the call and makes the agent keep two things in step. There is another way to build it.
Voicing Team9 min read
Real-Time TranslationA contact centre runs a successful three-month translation pilot. Agent feedback is positive. Handle times look reasonable. The business case holds.
Then it goes to IT for rollout planning, and someone asks how many desktops need the client installed.
The answer is four thousand, across six sites, on three operating system versions, with a browser plugin that needs to be maintained against Chrome’s release cycle. Each install needs security review. Each site needs a deployment window. The plugin needs to be re-certified after every browser update. Somebody has to own that, indefinitely.
The pilot was never the hard part. The pilot ran on twelve laptops that someone configured by hand.
This is where translation projects die, and the cause is not the translation. It is an architectural decision made long before anyone wrote a line of speech-processing code: the decision to put translation beside the call rather than inside it.
What “beside the call” actually means
Almost all translation tooling works the same way, whatever the interface looks like.
The call happens between customer and agent. In parallel, audio is captured and shipped to a separate system. That system transcribes, translates and returns a result, which is presented to the agent through a second surface, a desktop application, a browser extension, a panel inside the agent client, or a separate dial-in bridge.
The call and the translation are two distinct things happening simultaneously, and something has to keep them synchronised in real time.
That something is either a fragile integration or, more commonly, the agent.
Everything painful about bolt-on translation follows from this parallelism.
The agent becomes the integration layer. They hold the customer conversation in one window and the translation in another, reconciling them continuously while also doing their actual job. Cognitive load rises on precisely the calls that already demand the most.
Context fragments. The customer record lives in the CRM. The translation lives in the translation tool. Nothing connects them automatically, so the agent stitches two systems together under time pressure.
If the translation sits beside the call, the person holding the two together is the agent, on every single turn.
Every update is a training event. A second tool has its own release cycle, its own interface changes, its own quirks. Each change requires communication, and often retraining, across the entire agent population.
IT inherits a fleet-wide obligation. An application on every endpoint is something to deploy, secure, patch, audit and certify, multiplied by every desktop and repeated with every OS and browser update.
None of these costs are inherent to translation. They are inherent to translation happening somewhere other than the call.
Translation as a property of the connection
The alternative is to put the translation engine inside the call itself, at the signalling and media layer, so there is no second system to synchronise because there is no second system.
In SIP terms this is a back-to-back user agent, or B2BUA.
A B2BUA terminates the inbound call leg and originates a second leg toward the destination, bridging the two. It is not observing the call from outside. It is a participant in the connection, with native access to the media stream on either side.
That access is what makes real-time translation possible without a parallel system. Audio passing through the bridge is processed in flight: the customer’s speech rendered into the agent’s language on one leg, the agent’s speech into the customer’s on the other. Each direction is handled independently, which gives per-stream control over latency, quality and recording.
The media never leaves the path it was already on. There is no copy of the call routed elsewhere and reconciled afterward. The translation is not attached to the call. It is the call.
From the agent desktop’s perspective, nothing unusual is occurring. There is a phone call in progress. That ordinariness is the entire design objective, because it is what removes every cost listed in the previous section.
The correlation problem, and how signalling solves it
Bridging a call into two legs raises an immediate practical question, and it is the one that catches naive implementations.
If a call arrives, gets transferred to a translation platform, and then a new call is originated back toward the agent, the contact centre platform sees two separate conversations. It assigns the return leg its own conversation identifier. As far as reporting, recording and routing are concerned, these are unrelated calls.
That breaks everything downstream. The agent gets a call with no customer context. Reporting shows double the volume. Recording produces two unlinked audio files with no indication they are the same interaction.
The solution sits in the signalling rather than the media. SIP allows user-to-user information to travel alongside a call, and that channel carries the metadata needed to correlate the legs.
In our implementation, the contact centre platform generates a translation correlation identifier before transferring to the platform. The return leg carries the same identifier. A session and correlation manager maps both conversations to that single identifier, maintains session state, and emits call detail records and quality metrics against it.
The consequences are worth stating explicitly, because they are what makes the architecture deployable rather than merely elegant:
Context arrives with the call. The translated call reaches the agent already associated with the correct customer record, queue and interaction history. Context is preserved across the bridge rather than reconstructed after it.
Recording stays native. Each leg is recorded by the contact centre platform as it normally would be, then reconciled through the correlation identifier for unified reporting. You are not replacing your recording infrastructure or introducing a second archive.
Reporting reconciles. Volume, handle time and outcome data tie back to one interaction rather than appearing as two.
One design note that matters in evaluation: the legs are linked by a correlation identifier, not a shared conversation identifier. Contact centre platforms assign a new conversation ID to the return leg and there is no guarantee they are related. Any vendor claiming the two legs share a native conversation ID either has not built this or has misunderstood how the platform behaves.
Why the trunk is the right integration point
Because the integration happens at standard SIP trunking rather than inside a proprietary application surface, the architecture is not tied to any single contact centre platform.
Any operation that can present a SIP trunk, effectively any modern contact centre, can adopt it through bring-your-own-carrier configuration. The platform inserts into the trunk. Everything else stays where it is.
Three consequences follow.
No platform migration. The operation does not move to a new contact centre platform to gain translation. The capability meets the infrastructure where it already lives. This matters more than it sounds: platform migration is a multi-quarter programme with its own risk profile, and making translation contingent on one is how translation gets deferred indefinitely.
No re-architecture. Existing routing, queueing, reporting and agent tooling remain in place. Translation is added beneath them.
Portability. Because the integration point is the SIP layer rather than a vendor-specific application API, the same approach moves across platforms with configuration changes rather than rebuilds. An operation running two platforms after an acquisition does not need two translation implementations.
There is a broader principle here. A network-layer capability integrates with the network, which is why one architecture can serve many platforms without creating lock-in. Capabilities implemented at the application layer inherit the application’s boundaries.
What this does for security review
The security posture of in-path translation is different in kind, not just degree, from the bolt-on model.
With desktop tooling, audio is processed across many distributed client installations. The trust boundary is every endpoint. Securing it means securing thousands of machines, each capable of handling customer voice data, each a separate audit surface.
With in-path translation, audio is processed inside a single controlled media path. The trust boundary is one place, and it is a place designed to be governed.
Concretely:
One media path to secure rather than a fleet of endpoints. Signalling and media are secured with TLS and SRTP.
Deployable inside the enterprise perimeter. The capability can run in the customer’s own cloud tenancy or fully on-premise, so voice data stays inside the boundary the enterprise already controls and has already certified.
No endpoint software to certify. There is no client to patch against OS updates, no plugin to re-certify against browser releases, and no per-desktop audit obligation.
Complete and consistent records. Because the platform sits in the call, both original and translated audio can be captured, transcribed and logged consistently, producing a reconstructable record of what was said in each language. We’ve written about what auditability actually requires in a two-language conversation in how do you audit a conversation that happened in two languages?.
The practical effect in procurement is that the security review is a review of one system in a controlled environment, rather than a fleet-wide endpoint programme. That difference frequently determines whether a project gets approved at all.
Scaling behaviour
Two properties matter operationally.
Leg handling is stateless, which allows the B2BUA layer to scale horizontally with concurrent call volume rather than requiring vertical scaling of a stateful bridge.
And per-hop latency is attributed. Because the platform owns both legs and every processing stage between them, latency can be measured per stage and returned rather than inferred. A slow call is diagnosable from its own telemetry rather than requiring a profiler and a reproduction. Observability also covers hallucination and intent-drift detection across interactions, with explainable logging.
That attribution matters more than it initially appears. In a bolt-on architecture spanning a capture agent, a network hop, an external service and a return path, “the translation felt slow” is genuinely hard to diagnose. When one system owns the whole path, it is a lookup.
What to ask a vendor
Five questions separate in-path architectures from bolt-on tools quickly.
What gets installed on the agent desktop? If the answer is anything at all, the parallelism problem and all its costs are present.
How are the two call legs correlated? A specific answer involving signalling metadata indicates a real implementation. Vagueness, or a claim of a shared conversation ID, does not.
Does existing recording continue to work natively? If the vendor replaces your recording or introduces a second archive, your compliance position has changed and someone will need to re-certify it.
Which platforms are supported, and what is the integration point? Trunk-level integration means platform-agnostic. Application-level integration means a separate build per platform and a dependency on each vendor’s API surface.
Can it run inside our infrastructure boundary? For regulated operations this frequently decides the outcome regardless of any other answer.
The reframe
The translation market talks about models, languages and latency. Those matter, but they are not what determines whether a translation capability reaches production.
What determines it is a decision made at the architecture level about where translation happens. Beside the call, and you inherit a desktop application, a fleet deployment programme, a per-endpoint audit surface, and an agent doing synchronisation work by hand. Inside the call, and none of those exist to be solved.
If your multilingual strategy currently depends on software running on the agent’s screen, the architecture is working against the outcome you want.
- What is a SIP B2BUA in real-time translation?
- A back-to-back user agent terminates the inbound call leg and originates a second leg toward the destination, bridging the two. Because it participates in the connection rather than observing from outside, it has native access to the media stream on each side and can translate audio in flight.
- Why does bolt-on translation software create problems?
- Bolt-on tools sit parallel to the call, so the call and the translation are separate things that must be kept synchronised in real time, usually by the agent. This produces cognitive load, fragmented context, retraining with every tool update, and a fleet-wide deployment and audit obligation for IT.
- How are two call legs linked when translation bridges a call?
- Through metadata carried in the SIP signalling. A correlation identifier generated before the transfer travels with both legs, allowing a session manager to map both platform conversations to one interaction. Contact centre platforms assign a new conversation ID to the return leg, so a correlation identifier is required rather than a shared conversation ID.
- Does in-path translation require replacing my contact centre platform?
- No. Because the integration point is standard SIP trunking via bring-your-own-carrier configuration, existing routing, queueing, reporting and agent tooling stay in place. Translation is added beneath them at the trunk level.
- Does call recording still work with SIP-native translation?
- Yes. Each leg is recorded natively by the contact centre platform and reconciled through the correlation identifier for unified reporting, so existing recording infrastructure and retention policies remain in place.
- Is in-path translation easier to secure than desktop translation software?
- Generally yes, because the trust boundary is a single controlled media path rather than thousands of endpoints. There is no client software to patch, no plugin to re-certify against browser updates, and no per-desktop audit surface. The capability can also run inside the enterprise’s own infrastructure boundary.
Bring one call type. Leave with an architecture.

Voice infrastructure on the contact centre floor
A working session with an engineer who has deployed inside a bank’s perimeter. We map your telephony, data boundary and handoff rules, and tell you what we would not automate.