Field test · Call centre agent

Voice agents, measured on the line.

Four ways to put an agent on a phone call — one managed platform and three OpenAI architectures — running the same prompt, the same flow and the same scripted call, compared on cost and on the silence between turns.

Emiliano Merlo September 23, 2026
Figure 0 · The thing you hear first The silence before each answer
Typical
ElevenLabsmanaged pipeline
1.8s1.8–2.5s
Duplexgpt-live-1 + nano
2.1s2.1–2.3s
Realtime minispeech to speech
2.5s2.5–3.0s
Cascadethree models
5.2s5.2–7.4s
Customer speaking Agent speaking Silence · the agent is working
All four tracks share one time scale, zeroed on the customer's last word. Typical wait shown solid; the range spans the slowest 10% of answers. Cascade runs three models in sequence, and the wait is roughly three times everyone else's.

01 Introduction

With the emergence of agents, a new question came into the equation: how do we communicate with them? Voice is the most natural way. Voice agents have a variety of use cases that can make everyday tasks easier. In this post, we explore some options for a call center agent whose goals are to renew, upgrade or save a customer who wants to cancel.

For this purpose, we tried two approaches. The first is a full platform, ElevenLabs, which has flows, tools, test suites, post-call analysis, a large voice library and phone calls built in. The second is OpenAI, which offers several different ways of implementing it:

  • Realtime: one model that goes straight from speech to speech.
  • Cascade: three separate models for transcription, reasoning and speech.
  • Duplex: gpt-live-1 handles the voice communication, with a backend agent to run tools and do the thinking.

To test them, we created an agent for a fictional telecom company, running the same prompt and flow on all four. And since we wanted to test them on real phone calls, we connected the agents through Twilio.

Figure 1 · Architecture What sits between the customer and the answer
ElevenLabs
customer speech
▼
Scribe v2 Realtimespeech to text
▼
gemini-2.5-flashreasoning
▼
Eleven Flash v2text to speech
▼
agent speech
Realtime
customer speech
▼
gpt-realtime-2.1-minihears and speaks in one model
▼
agent speech
Cascade
customer speech
▼
gpt-4o-mini-transcribespeech to text
▼
gpt-4o-minireasoning
▼
gpt-4o-mini-ttstext to speech
▼
agent speech
Duplex
customer speech
▼
gpt-live-1voice, turns and interruptions
delegates▼
gpt-5.4-nanoprocedure, catalog and tools
▼
agent speech
The same conversation, four arrangements. Only ElevenLabs and the cascade put a text transcript in the middle; Realtime and Duplex keep the customer's audio in the model.

Each option was built with its platform's own recommended pattern and built-in tools: ElevenLabs' workflows, OpenAI's Agents SDK for Realtime and the cascade, and OpenAI's delegation guide for Duplex. We then ran the same scripted call on each, a customer who interrupts the offer, refuses it and asks to cancel, and measured cost and latency on those calls. They are small tests, one call per option, so take the numbers as indications rather than benchmarks.

Figure 2 · Scenario The call flow, and the path our test call took
Verificationconfirms it's the customer
it's them
UpsellTotal 100 · $27/month
declines
Renewalsame plan · $27/month
wants to cancel
RetentionEssential 70 · $17/month
accepts
Closeconfirms the plan, hangs up
Verification · wrong person or no interestSign off — apologizes, hangs up
Upsell · acceptsClose
Upsell · wants to cancel right awayRetention
Renewal · agrees to renewClose
Retention · still cancelsCancellation — processes it, hangs up
The blue path is the one our scripted customer took on every option: verified, turned down the upgrade, asked to cancel, then accepted the retention offer. The remaining branches were available to all four agents.

02 Listen to a call

Numbers only go so far — here are two real calls, recorded on the test line with gpt-realtime-2.1-mini answering. Each one is two-channel: the customer on one track, the agent on the other, never overlapping. Press play and watch where the time actually goes.

Roughly half of each call is neither side talking. Some of that is ordinary turn-taking, and some of it is the agent working: the longer gaps are where it went away to look something up before it could answer. The gap markers on the timeline are the real silence in the recording, whatever caused it.

03 How they work

ElevenLabs

Scribe v2 Realtimegemini-2.5-flashEleven Flash v2

ElevenLabs is a pipeline: it turns the customer's speech into text, decides what to answer, and turns that answer back into speech. All the models behind it are managed by the platform, so we never have to touch the audio ourselves and the setup feels seamless. Its workflows let you design the conversation as a graph, where each step has a clear goal and the model decides when to move to the next one. Each step can use tools and MCP integrations, including built-in tools that are very convenient for phone calls, like hanging up or transferring. Tools can also be given to specific steps only, so, for example, the agent can only hang up once it reaches a closing step.

What makes it stand out is everything around the agent. You can create a test suite to validate the different flows before going live, which is very valuable when working with LLM agents. After each call, data extraction gives you an analysis of the conversation. On the call itself, interruptions are built in and worked very well: the customer can cut the agent off at any moment and the conversation feels natural. The only weak spot we found was the recordings, where the audio sync is sometimes a bit off.

Figure 3 · ElevenLabs A managed pipeline, under two seconds before each answer
STT
LLM
TTS
STT
LLM
TTS
~1.8 s of silence
An example turn: the customer finishes, three managed models run in sequence, and the answer starts. Interruptions are handled by the platform.

OpenAI Realtime

gpt-realtime-2.1-mini

The Realtime API takes a different path: instead of a pipeline, a single model listens to the customer and answers directly in speech. This makes the connection very direct, with no intermediate steps between hearing and speaking. The trade-off is that there's no transcript of what the customer says unless another model transcribes it in parallel.

Since it's a single agent, the whole flow lives in its prompt. It supports tools and MCP, but there's no flow framework to handle the different situations in a more deterministic way, so every behaviour depends on how well the prompt is written. On the call itself, interruptions are also built in: the model detects when the customer starts talking and stops its answer, so the conversation sounds natural, with good intonation and quick replies.

On a phone line there's one extra step: the phone provider queues the agent's audio, so the application has to tell the model how much of it the customer actually heard. The Agents SDK includes a playback tracker for exactly that, and with it Realtime was the quickest of the four to stop when interrupted.

Figure 4 · Realtime One model from speech to speech, about 2.5 seconds before each answer
gpt-realtime-mini
gpt-realtime-mini
~2.5 s of silence
No transcription or speech step to wait on — the wait is the model itself deciding what to say.

OpenAI Cascade

gpt-4o-mini-transcribegpt-4o-minigpt-4o-mini-tts

The cascade follows the same architecture as ElevenLabs, but built by ourselves: one model transcribes what the customer says, another decides what to answer, and a third turns that answer into speech. Because the reply is plain text, this is the option that gives the most control, since the flow can be enforced in code instead of only requested in a prompt.

That control comes at a cost. OpenAI's Agents SDK covers part of what ElevenLabs gives you with its VoicePipeline: it detects when the customer has finished talking, runs the agent and streams the spoken reply. What it doesn't have is interruptions, so the agent always finishes its sentence, even when the customer talks over it. The bigger limitation is latency. Each reply goes through three models in sequence and in our test call the customer waited about five seconds for a typical answer, so the pauses between turns are too long for a natural conversation.

Figure 5 · Cascade Three models in sequence, about five seconds before each answer
STT
gpt-4o-mini
TTS
STT
gpt-4o-mini
TTS
~5.5 s of silence
Three round trips stack up before a single word is spoken — and the pipeline has no way to stop once it starts.

OpenAI Duplex

gpt-live-1gpt-5.4-nano

Duplex splits the work between two models. gpt-live-1 listens and speaks at the same time, just like a person on a call, handling the voice, the turns and the interruptions by itself. When it needs to think or run a tool, it delegates to a smaller backend model, gpt-5.4-nano in our case.

This makes it the option that sounds the most human. Since the model keeps listening while it speaks, interruptions are native, and it reacts while you talk instead of waiting for its turn. It's also the one with the least control: the voice model decides when to delegate and tools only exist on the backend side, so the flow depends on both models working well together.

That makes the prompts matter more than anywhere else. OpenAI's guide recommends splitting them: the backend gets the full procedure, the catalog and the tools, while the voice model only gets short rules on when to delegate. Those rules have to be explicit, including not guessing or filling the silence while the backend works, or the voice model improvises.

Figure 6 · Duplex A voice model delegating to a backend, about two seconds before each answer
gpt-live-1gpt-5.4-nano
gpt-live-1gpt-5.4-nano
~2 s of silence
The voice model holds the line while the backend thinks. Two stacked labels mean both models are working inside the same pause.

04 Features comparison

Beyond how each one talks, they differ in what they hand you and what you have to build. ElevenLabs is the only one where the flow is enforced by the platform instead of asked for in a prompt, and the only one with a knowledge base ready to use. On the OpenAI options the flow is prompt-scoped, except in the cascade, where the reply is text before it is spoken, so the orchestration is your own code and can be as strict as you like.

Figure 7 · Capabilities What each option gives you
Flow controlKnowledgeToolsMCP
ElevenLabs Visual graph, enforced by the platform Managed, with RAG Built-in and custom, scoped per step External MCP servers
Realtime In the prompt; instructions and tools can change mid-call In the prompt, or fetched by a tool Function tools, declared in the session format MCP servers on the session, run by the API
Cascade Your own code: any orchestration, replies checked before speaking Whatever your code calls Whatever your code calls Whatever your code calls
Duplex In the voice prompt; the procedure lives in the backend On the backend agent On the backend agent On the backend agent
Flow, knowledge and tools across the four options. Where a cell reads whatever your code calls, nothing ships with the platform — the capability is yours to build, and yours to constrain.

05 Price

Prices exclude Twilio, which costs the same per minute for every option. Figures come from the test calls, divided by the call length Twilio reported. Keep in mind, though, that Twilio charges by call duration: an agent that takes longer to answer makes the same conversation last longer, so a slower option also adds more telephony cost on top of its own price.

Figure 8 · Cost Price per minute
ElevenLabs
$0.080
Realtime 2.1full
~$0.080
Duplex
$0.050 + backend
Realtime 2.1mini
$0.021
Cascade
$0.011 *
$0.00$0.02$0.04$0.06$0.08
Hatched — estimated by applying gpt-realtime-2.1 rates to our Realtime mini call. Dashed — Duplex's backend model is billed separately and not reported. * Transcription and speech priced from duration.
USD per minute, measured on our test calls, Twilio excluded.

The full Realtime model is about as expensive as ElevenLabs, roughly 4 times its mini version. On the other end, ElevenLabs costs about 4 times as much as Realtime mini and around 7 times as much as the cascade. A few things to keep in mind:

  • ElevenLabs: the price depends on your plan.
  • Duplex: its price is a floor, because the backend model is billed separately.
  • Realtime: every turn re-sends the conversation's audio, so longer calls cost more per minute. Cached audio is billed at a small fraction of the normal price, which keeps that growth under control.
  • Cascade: most of its cost is speech generation, not the reasoning model: about two thirds in our test call.

06 Latency

We measured latency on each call's recording, which keeps the customer and the agent on separate channels. It's the silence between the customer's last word and the agent's first.

Figure 9 · Latency Latency on our test calls
Waiting for an answer
ElevenLabs
1.8s – 2.5s
Duplex
2.1s – 2.3s
Realtime mini
2.5s – 3.0s
Cascade
5.2s – 7.4s
0s2s4s6s8s
Kept talking after being interrupted
Realtime mini
0.5s
ElevenLabs
1.2s
Duplex
1.6s
Cascade
Doesn't stop
0s0.5s1s1.5s2s
Measured from each call's recording, one call per option. Solid is the typical wait; faded is the slowest 10% of answers. Shorter is better in both panels.
  • ElevenLabs answered fastest.
  • Duplex was the most consistent: its slowest answers were barely slower than a typical one.
  • Realtime was the quickest to stop talking when the customer cut in.
  • Cascade made the customer wait around five seconds, and its pipeline doesn't support interruptions.

07 Conclusion

There's no single winner; it depends on what you prioritize.

All four agents were natural and engaging in the conversation: the voices sounded good and the tone stayed warm and friendly. What matters most for it to feel like a real conversation is latency: the pauses between turns have to be short and kept under control, or the call quickly stops flowing. Beyond that, the options differ in speed, control and cost:

Fastest, fullest, priciest

ElevenLabs

If you want to go to production fast with good tooling, it gives you flows, tests, analysis and voices out of the box, and it answered fastest in our tests, at the highest price.

Wait
1.8s
Price
$0.080/min
The best balance

Realtime mini

It's cheap and natural, and the quickest to stop when interrupted, but the flow lives only in the prompt and you build the platform around it.

Wait
2.5s
Price
$0.021/min
Most control, lowest cost

Cascade

Good for strict scripts, if you can live with the latency of around five seconds per answer and no interruptions.

Wait
5.2s
Price
$0.011/min
Most human, steadiest

Duplex

The most human conversation and the steadiest latency. It's promising, but it's not the cheapest, and its flow is the hardest to control: both the voice model and the backend need careful prompts.

Wait
2.1s
Price
$0.050/min +

Latency isn't the only thing that makes a call feel real, though. Listening matters too, and we didn't evaluate it: our calls were scripted and quiet, so we didn't test whether an agent waits while the customer thinks, stays silent when asked, or handles background noise like traffic or other voices. These matter a lot on real calls, and they are worth testing before choosing one.

Related posts