Decagon Voice 3: what Chord and the duplex architecture actually change

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 4, 2026

Expert Verified
Hand-drawn illustration of a person at a laptop talking with an AI agent wearing a headset, with a sound wave between them and the Decagon logo on a purple band

What Decagon actually launched

Decagon shipped four things on October 1, all at once. Voice 3 grabbed the headlines. It makes more sense, though, when you see it next to the other three.

Hand-drawn map of the four Decagon Dialogues October 2026 launches: Voice 3 (Chord plus duplex) highlighted, Personal Agent Gateway, Agent Modules and Duet Apprentice
Hand-drawn map of the four Decagon Dialogues October 2026 launches: Voice 3 (Chord plus duplex) highlighted, Personal Agent Gateway, Agent Modules and Duet Apprentice

Per the Dialogues 2026 announcement from co-founders Jesse Zhang and Ashwin Sreenivas, the four releases are:

  • Voice 3: Chord plus a duplex architecture, for more natural phone calls.
  • Personal Agent Gateway: detection and a dedicated channel for personal AI agents (Decagon names Instinct, Muse and dots) that call or message your business on a customer's behalf.
  • Agent Modules: a base for stretching the agent past support into lead qualification, onboarding and collections.
  • Duet Apprentice: Duet learns from your internal docs and escalated conversations, then follows Slack and Microsoft Teams threads to draft procedure updates.

I build AI agents at eesel, and what jumped out to me is how much of this launch is about the phone. Decagon's own framing is that voice is "the least forgiving" channel, and Voice 3 is basically their answer to it. If you're new to the company itself, my Decagon overview and the Decagon review cover the wider platform.

The Decagon Voice product page with the headline "A voice customers will thank" and a demo video thumbnail, as taken from Decagon
The Decagon Voice product page with the headline "A voice customers will thank" and a demo video thumbnail, as taken from Decagon

How the duplex architecture works

If I were evaluating Voice 3, this is where I'd spend most of my time. It changes how a call with an AI agent feels.

Most AI voice agents run what Decagon calls a cascaded pipeline. Speech-to-text transcribes the caller and an LLM decides what to say, then text-to-speech reads the answer back. The Voice 3 announcement spells out the problem: each stage adds delay, a simple "mhm" can trigger an interruption, and a long lookup leaves the caller listening to silence.

Hand-drawn comparison of a cascaded voice pipeline (speech to text, LLM decides, text to speech, caller waits in silence) and the Voice 3 duplex design with two parallel lanes: talks and listens, and reasons, calls tools, guardrails
Hand-drawn comparison of a cascaded voice pipeline (speech to text, LLM decides, text to speech, caller waits in silence) and the Voice 3 duplex design with two parallel lanes: talks and listens, and reasons, calls tools, guardrails

Voice 3 splits the job into two layers that run at the same time:

  1. A low-latency conversational model does the listening and speaking, including quick progress updates.
  2. A more powerful model handles the reasoning, tool calls and guardrail checks behind the conversation.

So the agent keeps processing incoming audio while it's still speaking. Per the Dialogues post, it talks through a caller saying "mhm" but yields on a genuine interruption, and it can narrate progress or answer a follow-up while a long task runs instead of putting the caller on hold.

A quick example of where this matters. A customer calls about a double charge. A cascaded agent says "let me look that up," then goes quiet for six seconds while it hits the billing API. The caller says "hello?", which the agent reads as an interruption, and the whole thing restarts. A duplex agent can say "I'm pulling up your last two invoices now" and answer "is it the one from Tuesday?" while the lookup finishes. It's the gap between giving a bot one command at a time and actually having a conversation with it.

Decagon's voice product page lists the rest of the toolkit around it: interruption handling, hundreds of voice profiles including custom-tuned ones, outbound calling, transfers with a summary, and guardrails tied to its Agent Operating Procedures. If you want the broader picture of how call center automation fits together, that post is a decent primer.

What Chord is, and how Decagon built it

Chord is the first voice model from Decagon Labs. The Chord deep dive from Samuel Zhang explains the bet: general speech models are built for everything from audiobook narration to game voiceovers, and that caps how natural they sound on a live support call.

A few of the choices stand out:

  • Messy training data, on purpose. Most speech pipelines scrub out pauses and fillers. Decagon kept them, using verbatim transcription that preserves pacing, emphasis and disfluencies.
  • A recording method that isn't a script read. Voice actors handed a script give a performance, so Decagon built a method that mixes script content with free-flowing conversation.
  • A tokenizer-free diffusion base model. Most speech models chop audio into tokens and lose detail at each step. Decagon says skipping tokenization keeps Chord close to lossless and makes it easier to steer emotion, pace and emphasis.

In practice, Chord shapes speech phrase by phrase. It slows down for a confirmation code or phone number, then returns to a normal pace, instead of running on one global speed setting. Decagon says it's served on Modal in production and trained on licensed data and consented voice talent, never customer-owned data.

Flag that last point to your security team early. If you've ever been through a voice-cloning review, you know it's the first question they'll ask. It's also one of the few vendor statements here you can paste straight into a procurement doc.

Reading the Voice 3 numbers honestly

Decagon published three sets of results. All three come from Decagon itself. So yes, vendor data. Still, they're more specific than what most launch posts give you, and worth a close read.

Hand-drawn scorecards: 45.7% picked Chord as the human (coin flip is 50%), 36.2% preferred Chord (185 listeners, 4 voices), and plus 2.6 to plus 7.1 resolution points on live calls (3 customers, before vs after)
Hand-drawn scorecards: 45.7% picked Chord as the human (coin flip is 50%), 36.2% preferred Chord (185 listeners, 4 voices), and plus 2.6 to plus 7.1 resolution points on live calls (3 customers, before vs after)

The "90% couldn't tell" claim

The launch post says roughly 90% of listeners couldn't tell Chord from a real person. The chart underneath tells you more than the headline does. In a blind test across three voices, 54.3% picked the real human and 45.7% picked Chord as the human.

Decagon's blind listening test chart: 54.3% picked the human and 45.7% picked Chord, with a dashed line marking 50% chance, as taken from Decagon
Decagon's blind listening test chart: 54.3% picked the human and 45.7% picked Chord, with a dashed line marking 50% chance, as taken from Decagon

A perfect coin flip would be 50/50, so 45.7% sits close to "indistinguishable." The "90%" framing is roughly 45.7 divided by 50. Neither reading is wrong. Internally, I'd quote the 45.7% figure, since that's the one your team can actually check.

The preference test

Decagon also ran Chord against three other leading speech models with the same words and asked which voice listeners would rather hear on a support call. Across 185 listeners, Chord came first at 36.2%, ahead of 26.8%, 19.9% and 17.2% for the other three. Decagon doesn't name the competitors, so you can't map this onto ElevenLabs or anyone else directly.

Decagon's preference chart: Chord at 36.2% versus three unnamed speech models at 26.8%, 19.9% and 17.2%, n=185, as taken from Decagon
Decagon's preference chart: Chord at 36.2% versus three unnamed speech models at 26.8%, 19.9% and 17.2%, n=185, as taken from Decagon

The live-call results (the ones that matter)

As a buyer, this is the test I'd care about most. Decagon took three customers in telecom, financial services and travel who moved from an off-the-shelf voice to Chord, and compared the same programs before and after, changing only the voice.

IndustryResolution rate changeBarge rate change
Telecom+6.1 points-14.5 points
Financial services+7.1 points-6.1 points
Travel and hospitality+2.6 points-5.3 points

Source: Decagon's Chord post, October 1, 2026.

"Barge rate" is the share of calls where the caller only wants a human and keeps saying "representative" until they get one. A 14.5-point drop in telecom is a lot for a change that only touched the voice.

Decagon's bar charts of Chord's absolute lift in resolution rate (6.1, 7.1, 2.6 points) and reduction in barge rate (14.5, 6.1, 5.3 points) across telecom, financial services and travel, as taken from Decagon
Decagon's bar charts of Chord's absolute lift in resolution rate (6.1, 7.1, 2.6 points) and reduction in barge rate (14.5, 6.1, 5.3 points) across telecom, financial services and travel, as taken from Decagon

Two caveats, though. It's three customers, and Decagon doesn't publish the starting resolution rates, so you can't tell whether +7.1 means 50% to 57% or 80% to 87%. And the travel lift is a third the size of the financial services one, so the voice clearly matters more for some call types than for others.

The voice isn't the whole story (Decagon says so too)

The most useful bit, for me, comes from Decagon's own blog. In an earlier post on reducing barge, Decagon says it cut barge by over 15% for a global telecom on billing disputes "not with a new model or a better voice," but by changing what the agent said in its first few turns. That was one review of real conversations, shipped in about a day.

So the same company that just built a custom speech model is telling you the script can move barge as much as the voice. I think that's the honest way to read Voice 3. Chord gets callers to listen and the duplex design keeps them on the line. Whether the call actually gets resolved still comes down to the procedures behind it.

I learned the same lesson on the text side. At eesel, I've watched a confident-sounding AI agent give a polished, wrong answer more than once, which is why every rollout I help with gets simulated against past tickets before it goes live. A natural voice makes a wrong answer more convincing, not less. Whatever phone agent you pick, test it on your own call recordings, not on a vendor's blind test.

For context on where voice AI sits overall, the top score on Artificial Analysis' tau-Voice benchmark, which scores simulated support calls on whether the task actually gets done, was 68.6% at the end of September 2026. So even the best voice models still miss roughly a third of realistic support calls. My guide to AI phone support and the question of whether AI can answer support calls go deeper on that gap.

Languages, pricing and who it's for

Three practical questions decide whether Voice 3 is worth a demo: does it speak your customers' languages, what does it cost, and does your team look like Decagon's customers.

70+ languages from one agent

Per the Voice 3 post, Decagon supports 70+ languages without a separate agent per language. It detects the caller's language and switches automatically, even mid-sentence, uses locale-specific voices, and validates each language with native speakers before it ships. Deutsche Telekom is the named customer on the launch, and its digital service lead puts the shift plainly:

"Customers have been trained for years to speak in short fragments at these IVRs just to get through to a human."

Christian Niedworok, Lead of Digital Service Communication at Deutsche Telekom, via Decagon

Pricing

Decagon doesn't publish a price for Voice 3, and decagon.ai/pricing returns a 404. Every call to action is "Get a demo." Decagon's own glossary entry on resolution-based pricing describes charging when the AI completes a task without human help, which is the model buyers report, but no rate card exists. My Decagon pricing post collects what's actually known, and the Sierra pricing and Parloa pricing breakdowns show the same quote-only pattern across the enterprise tier.

Who should look at Voice 3

You are...My take
An enterprise with heavy, multilingual phone volumeShortlist it. Chord and duplex target your exact pain, and Decagon's named voice customers (Deutsche Telekom, Chime) look like you.
Already on Decagon for chat or emailAsk for the live-call numbers on your own traffic. Cross-channel memory is the quiet benefit.
A mid-size team that wants to build its own phone agentLook at Retell AI or the other platforms in my voice agent roundup, which publish per-minute pricing.
A team whose support is mostly email and chatVoice 3 isn't your first problem. Put AI on the ticket queue first.

If you're comparing Decagon against its peers, the Decagon vs Sierra and Decagon vs Zendesk AI posts, plus my list of Decagon alternatives, are the natural next reads. Zendesk shops should also look at Zendesk voice AI agents before signing a separate phone contract.

What happens after the call

Every phone agent hands some calls off. Decagon's voice page says Voice 3 transfers calls to human agents "with a concise summary," and in most support teams that summary becomes a ticket, a callback, or an email follow-up. That's all written work. It piles up in the helpdesk no matter which voice the caller heard.

One eesel customer told the eesel team their agents had to summarize every voice interaction by hand before the AI could use it, and that fixing it would significantly increase their usage. That gap between the phone and the queue is worth planning for before you pick a voice vendor. The AI agent handoff best practices post walks through what a good transfer needs to carry.

eesel for the tickets your phone agent hands off

eesel is an AI teammate platform, and the teammate that fits here is the AI helpdesk one. It joins your existing helpdesk the way a new hire would, through the Zendesk integration or Freshdesk, Gorgias, Front and others, learns from your past tickets and help center, and drafts or sends replies on the follow-ups a phone agent creates. Before it goes live, you can run a simulation on hundreds of your past tickets and see how it would have answered.

eesel AI dashboard showing connected Zendesk ticket activity
eesel AI dashboard showing connected Zendesk ticket activity

To be plain about it, eesel doesn't answer the phone. Paired with Voice 3 or any other AI call center agent, it handles the written half of support. Plans start with 100 free credits on the eesel pricing page, no card needed. Try eesel on your own ticket history and see what it would have sent.

Frequently Asked Questions

What is Decagon Voice 3?
Decagon Voice 3 is the third generation of Decagon's phone agent, launched on October 1, 2026. It speaks through Chord, a speech model Decagon post-trained on customer conversations, and runs on a duplex architecture that lets the agent keep talking while it looks things up. It's part of the wider Decagon platform covered in my Decagon review.
What is Decagon's Chord speech model?
Chord is the first voice model from Decagon Labs. It's a tokenizer-free diffusion model post-trained on conversational data, so it keeps the pauses, emphasis and pacing changes of a real call instead of reading like a script. Decagon says it's trained on licensed data and consented voice talent, never on customer-owned data. For how it compares with general speech models, see my roundup of the best AI voice agents.
How much does Decagon Voice 3 cost?
Decagon doesn't publish Voice 3 pricing, and its pricing page returns a 404. Every path runs through a sales demo, and Decagon's own glossary describes per-resolution billing. My Decagon pricing breakdown covers what's actually known about contracts.
How many languages does Decagon Voice 3 support?
Decagon says Voice 3 supports 70+ languages from one agent, detects the caller's language automatically, and can switch even when someone changes language mid-sentence. Each language is validated by native speakers before it ships. If most of your multilingual volume is text, multilingual live chat is worth sorting out first.
Is Decagon Voice 3 better than Sierra or Parloa for phone support?
Decagon is the only one of the three shipping its own post-trained speech model, which is the Voice 3 headline. Sierra and Parloa are both enterprise, sales-led platforms too, and none of them publishes a self-serve price. The Decagon vs Sierra comparison goes deeper on the trade-offs.
Can a voice AI agent fully replace phone support agents?
Not yet. Even Decagon's best published Chord result lifted resolution by 7.1 points, which still leaves a share of calls going to people. Plan for a clean handoff and read the AI agent handoff best practices before you switch on any phone agent.
What happens to calls Decagon Voice 3 can't resolve?
Decagon's voice page says the agent transfers calls to a human with a concise summary. In most support teams that transfer turns into a ticket or a follow-up in the helpdesk, which is where an AI helpdesk agent like eesel can pick up the written side of the work.
Who should consider Decagon Voice 3?
Large enterprises with heavy phone volume, multilingual callers and the budget for a vendor-led rollout are the natural fit, which matches Decagon's named customers like Deutsche Telekom and Chime. Smaller teams usually get more from automating phone support upstream and putting AI on the ticket queue first.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Hand-drawn illustration of a caller on a phone whose voice flows into a row of four AI agent cards, above a call timer and waveform
Guides

The 8 best AI voice agents in 2026, compared by who runs them

The best AI voice agents in 2026, split by who builds and runs them, with real per-minute costs, fresh benchmark scores, and where each one fits.

Rama AdiRama AdiSep 29, 2026
Hand-drawn illustration of customers handing an AI agent key to a support rep at a door marked with the Decagon logo, holding a ring of permission keys
Trending

Decagon Personal Agent Gateway: how it handles AI agents that call your support

Decagon's Personal Agent Gateway detects customers' AI agents, gives them their own channel and scopes what they can do. Here's how it works and what's still unknown.

Rama AdiRama AdiOct 6, 2026
Line illustration of a developer and a support agent talking through voice waveforms, next to the Grok logo
Trending

Grok Voice Think Fast 2.0 pricing: what $0.08/min costs

Grok Voice Think Fast 2.0 is $0.08 per minute of audio, 60% above 1.0. The grok-voice-latest alias moves to it today, so here is the real per-call math.

Kurnia KharismaKurnia KharismaAug 5, 2026
Line illustration of a support agent on a headset next to the Grok logo, with a voice waveform in a speech bubble
Trending

Grok Voice Think Fast 2.0: what changed and what it costs

Grok Voice Think Fast 2.0 scores 82.9% on the Artificial Analysis speech-to-speech index and answers in 0.70s. It also costs 60% more per minute, and the default alias flips to it on August 5.

Rama AdiRama AdiAug 4, 2026
Wonderful AI review illustration showing an enterprise AI agent across channels
Trending

Wonderful AI review (2026): is the enterprise AI OS worth it?

An honest Wonderful AI review: what the enterprise AI OS does, the funding story, the quote-only pricing, and who it actually fits in 2026.

Riellvriany IndriawanRiellvriany IndriawanSep 8, 2026
Blue waveform icon on an abstract blue and purple background
Guides

OpenAI Audio Speech API: what to build and test in 2026

Understand the OpenAI Audio Speech API for text-to-speech, transcription, and realtime voice work, then separate the model infrastructure from the support workflow you operate.

Rama AdiRama AdiOct 12, 2025
Two people talking across a table while an audio-visual AI model watches, listens and speaks in the same loop
Trending

SeedRealtime: what ByteDance's audio-visual model actually does

SeedRealtime is ByteDance's audio-visual full-duplex model. Here is what it does, what ByteDance published, and what you can actually call today.

KiraKiraAug 18, 2026
Hand-drawn illustration of two people in a natural back-and-forth voice conversation, in OpenAI teal
Trending

GPT-Live-1 review: is OpenAI's full-duplex voice model worth it?

A hands-on GPT-Live-1 review: OpenAI's full-duplex voice model is the most natural I have used, but it delegates reasoning to a backend model and bills per minute. Here is the honest verdict.

KiraKiraSep 11, 2026
Two people having a natural conversation with an AI voice assistant, sound waves flowing between them
Trending

GPT-Live-1: OpenAI's full-duplex voice model, explained

What GPT-Live-1 actually is: OpenAI's full-duplex voice model that listens and speaks at once, now in ChatGPT and the API at $0.05 per minute.

KiraKiraSep 11, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free