01 Introduction
With the emergence of agents, a new question came into the equation: how do we communicate with them? Voice is the most natural way. Voice agents have a variety of use cases that can make everyday tasks easier. In this post, we explore some options for a call center agent whose goals are to renew, upgrade or save a customer who wants to cancel.
For this purpose, we tried two approaches. The first is a full platform, ElevenLabs, which has flows, tools, test suites, post-call analysis, a large voice library and phone calls built in. The second is OpenAI, which offers several different ways of implementing it:
- Realtime: one model that goes straight from speech to speech.
- Cascade: three separate models for transcription, reasoning and speech.
- Duplex: gpt-live-1 handles the voice communication, with a backend agent to run tools and do the thinking.
To test them, we created an agent for a fictional telecom company, running the same prompt and flow on all four. And since we wanted to test them on real phone calls, we connected the agents through Twilio.
Each option was built with its platform's own recommended pattern and built-in tools: ElevenLabs' workflows, OpenAI's Agents SDK for Realtime and the cascade, and OpenAI's delegation guide for Duplex. We then ran the same scripted call on each, a customer who interrupts the offer, refuses it and asks to cancel, and measured cost and latency on those calls. They are small tests, one call per option, so take the numbers as indications rather than benchmarks.
02 Listen to a call
Numbers only go so far — here are two real calls, recorded on the test line with gpt-realtime-2.1-mini answering. Each one is two-channel: the customer on one track, the agent on the other, never overlapping. Press play and watch where the time actually goes.
Roughly half of each call is neither side talking. Some of that is ordinary turn-taking, and some of it is the agent working: the longer gaps are where it went away to look something up before it could answer. The gap markers on the timeline are the real silence in the recording, whatever caused it.
03 How they work
ElevenLabs
ElevenLabs is a pipeline: it turns the customer's speech into text, decides what to answer, and turns that answer back into speech. All the models behind it are managed by the platform, so we never have to touch the audio ourselves and the setup feels seamless. Its workflows let you design the conversation as a graph, where each step has a clear goal and the model decides when to move to the next one. Each step can use tools and MCP integrations, including built-in tools that are very convenient for phone calls, like hanging up or transferring. Tools can also be given to specific steps only, so, for example, the agent can only hang up once it reaches a closing step.
What makes it stand out is everything around the agent. You can create a test suite to validate the different flows before going live, which is very valuable when working with LLM agents. After each call, data extraction gives you an analysis of the conversation. On the call itself, interruptions are built in and worked very well: the customer can cut the agent off at any moment and the conversation feels natural. The only weak spot we found was the recordings, where the audio sync is sometimes a bit off.
OpenAI Realtime
The Realtime API takes a different path: instead of a pipeline, a single model listens to the customer and answers directly in speech. This makes the connection very direct, with no intermediate steps between hearing and speaking. The trade-off is that there's no transcript of what the customer says unless another model transcribes it in parallel.
Since it's a single agent, the whole flow lives in its prompt. It supports tools and MCP, but there's no flow framework to handle the different situations in a more deterministic way, so every behaviour depends on how well the prompt is written. On the call itself, interruptions are also built in: the model detects when the customer starts talking and stops its answer, so the conversation sounds natural, with good intonation and quick replies.
On a phone line there's one extra step: the phone provider queues the agent's audio, so the application has to tell the model how much of it the customer actually heard. The Agents SDK includes a playback tracker for exactly that, and with it Realtime was the quickest of the four to stop when interrupted.
OpenAI Cascade
The cascade follows the same architecture as ElevenLabs, but built by ourselves: one model transcribes what the customer says, another decides what to answer, and a third turns that answer into speech. Because the reply is plain text, this is the option that gives the most control, since the flow can be enforced in code instead of only requested in a prompt.
That control comes at a cost. OpenAI's Agents SDK covers part of what ElevenLabs gives you with its VoicePipeline: it detects when the customer has finished talking, runs the agent and streams the spoken reply. What it doesn't have is interruptions, so the agent always finishes its sentence, even when the customer talks over it. The bigger limitation is latency. Each reply goes through three models in sequence and in our test call the customer waited about five seconds for a typical answer, so the pauses between turns are too long for a natural conversation.
OpenAI Duplex
Duplex splits the work between two models. gpt-live-1 listens and speaks at the same time, just like a person on a call, handling the voice, the turns and the interruptions by itself. When it needs to think or run a tool, it delegates to a smaller backend model, gpt-5.4-nano in our case.
This makes it the option that sounds the most human. Since the model keeps listening while it speaks, interruptions are native, and it reacts while you talk instead of waiting for its turn. It's also the one with the least control: the voice model decides when to delegate and tools only exist on the backend side, so the flow depends on both models working well together.
That makes the prompts matter more than anywhere else. OpenAI's guide recommends splitting them: the backend gets the full procedure, the catalog and the tools, while the voice model only gets short rules on when to delegate. Those rules have to be explicit, including not guessing or filling the silence while the backend works, or the voice model improvises.
04 Features comparison
Beyond how each one talks, they differ in what they hand you and what you have to build. ElevenLabs is the only one where the flow is enforced by the platform instead of asked for in a prompt, and the only one with a knowledge base ready to use. On the OpenAI options the flow is prompt-scoped, except in the cascade, where the reply is text before it is spoken, so the orchestration is your own code and can be as strict as you like.
| Flow control | Knowledge | Tools | MCP | |
|---|---|---|---|---|
| ElevenLabs | Visual graph, enforced by the platform | Managed, with RAG | Built-in and custom, scoped per step | External MCP servers |
| Realtime | In the prompt; instructions and tools can change mid-call | In the prompt, or fetched by a tool | Function tools, declared in the session format | MCP servers on the session, run by the API |
| Cascade | Your own code: any orchestration, replies checked before speaking | Whatever your code calls | Whatever your code calls | Whatever your code calls |
| Duplex | In the voice prompt; the procedure lives in the backend | On the backend agent | On the backend agent | On the backend agent |
05 Price
Prices exclude Twilio, which costs the same per minute for every option. Figures come from the test calls, divided by the call length Twilio reported. Keep in mind, though, that Twilio charges by call duration: an agent that takes longer to answer makes the same conversation last longer, so a slower option also adds more telephony cost on top of its own price.
The full Realtime model is about as expensive as ElevenLabs, roughly 4 times its mini version. On the other end, ElevenLabs costs about 4 times as much as Realtime mini and around 7 times as much as the cascade. A few things to keep in mind:
- ElevenLabs: the price depends on your plan.
- Duplex: its price is a floor, because the backend model is billed separately.
- Realtime: every turn re-sends the conversation's audio, so longer calls cost more per minute. Cached audio is billed at a small fraction of the normal price, which keeps that growth under control.
- Cascade: most of its cost is speech generation, not the reasoning model: about two thirds in our test call.
06 Latency
We measured latency on each call's recording, which keeps the customer and the agent on separate channels. It's the silence between the customer's last word and the agent's first.
- ElevenLabs answered fastest.
- Duplex was the most consistent: its slowest answers were barely slower than a typical one.
- Realtime was the quickest to stop talking when the customer cut in.
- Cascade made the customer wait around five seconds, and its pipeline doesn't support interruptions.
07 Conclusion
There's no single winner; it depends on what you prioritize.
All four agents were natural and engaging in the conversation: the voices sounded good and the tone stayed warm and friendly. What matters most for it to feel like a real conversation is latency: the pauses between turns have to be short and kept under control, or the call quickly stops flowing. Beyond that, the options differ in speed, control and cost:
ElevenLabs
If you want to go to production fast with good tooling, it gives you flows, tests, analysis and voices out of the box, and it answered fastest in our tests, at the highest price.
- Wait
- 1.8s
- Price
- $0.080/min
Realtime mini
It's cheap and natural, and the quickest to stop when interrupted, but the flow lives only in the prompt and you build the platform around it.
- Wait
- 2.5s
- Price
- $0.021/min
Cascade
Good for strict scripts, if you can live with the latency of around five seconds per answer and no interruptions.
- Wait
- 5.2s
- Price
- $0.011/min
Duplex
The most human conversation and the steadiest latency. It's promising, but it's not the cheapest, and its flow is the hardest to control: both the voice model and the backend need careful prompts.
- Wait
- 2.1s
- Price
- $0.050/min +
Latency isn't the only thing that makes a call feel real, though. Listening matters too, and we didn't evaluate it: our calls were scripted and quiet, so we didn't test whether an agent waits while the customer thinks, stays silent when asked, or handles background noise like traffic or other voices. These matter a lot on real calls, and they are worth testing before choosing one.


