Sierra fleming-1: the model that tells you when an AI agent is calling

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 11, 2026

Expert Verified
Hand-drawn illustration of a support agent at a desk taking a call from a robot caller while a colleague talks to an AI assistant, with the Sierra logo on a green background

What is Sierra fleming-1?

fleming-1 is the newest model in what Sierra calls its "constellation of models", the set of fine-tuned models behind its Agent OS. Sierra, one of the bigger agentic customer service platforms, named it on stage at Sierra Summit 2026 alongside Curie, a model that runs the core loop of a conversation. The summit recap describes Fleming as the model that "detects when the caller on the phone is another agent, enabling businesses to handle bots and fraud as well as personal agents."

A dedicated post titled "Caller ID in the age of agents" followed, written by Ajeet Grewal and Venu Satuluri and dated October 8, 2026. Bret Taylor, Sierra's co-founder, shared it the evening before.

Sierra's announcement post for fleming-1, titled Caller ID in the age of agents, as taken from Sierra's blog
Sierra's announcement post for fleming-1, titled Caller ID in the age of agents, as taken from Sierra's blog

I've spent my time at eesel building the AI agents that sit in support queues, and the most useful thing about this launch is how narrow it is. fleming-1 does one job: it answers "is this voice synthetic?" while the call is still happening. It does not authenticate anyone, it does not judge intent, and it does not touch chat or email. Keep that scope in your head and the rest of the launch makes sense.

Here is everything Sierra has confirmed, in one place:

DetailWhat Sierra says
What it detectsCallers whose speech is likely AI-generated
ChannelPhone calls only
WhenIn real time, during the call
Default tuningConservative, so real people are not flagged by mistake
OutputA flag on calls it identifies as likely AI
What happens nextYour business decides
Who can use itAny voice agent built on Sierra
How to enable"You just need to turn it on"
PriceNot published
Accuracy, false positives, latency, languagesNot published

Why Sierra built a detector for AI callers

Sierra says it first hit the problem in 2025, "when agents built on Sierra started calling agents built on Sierra in healthcare." Then the consumer side took off. Personal agents such as Meta Muse and Instinct now wait on hold, work phone trees and talk to reps on their owners' behalf. Sierra's post warns that "a large share of the calls companies receive could come from AI acting on behalf of consumers."

The people on the other end of those calls often cannot tell. One researcher described what happened when Muse called a company for him:

"This week, Muse called customer service on my behalf. It navigated the phone tree, waited on hold, spoke with a human rep, and resolved the issue (worked great!). The crazy thing for me (besides that it works) is that the rep didn't blink. Talking to a bot was completely natural to them."

Some agents do say what they are. A Muse user on Reddit noticed their agent opening calls by naming who it worked for:

Reddit

"It says Hi I am Brett calling on behalf of gadgetneil and im taking notes etc."

That is the polite case. The worry is the agent that does not introduce itself, at scale. How big is the volume already? The closest public number comes from a fraud vendor, not Sierra: Pindrop says that in two months it identified 727,000 AI agents or bots on the calls it screens, roughly 1 in 80 calls, according to its BotStopper page. That is Pindrop's own customer base, not the whole market, but it is a useful order of magnitude.

The most-shared take on where this goes was blunt about the incentive problem:

"AI agents reduce the cost of complaints/requests to ~ zero. Anything free is consumed at much higher rates, so we should expect the total volume of this category to explode."

Sierra's answer is not to shut that traffic out, which fits how it pitches AI call center agents generally. Its post is careful to say some automated callers are fraudsters testing account security, while "others will be customers who handed off a chore to an agent, or people who rely on text-to-speech." That last group is the reason a blunt "block all AI voices" rule would be a mistake. Text-to-speech users are real customers.

How fleming-1 works

Sierra has shared the shape of the system, not the internals. fleming-1 "analyzes callers' speech in real time, scoring the audio for signs that it was generated by AI." Sierra says it looks "beyond how a voice sounds to a human listener", which matters because synthetic voices now fool people easily, and background noise or a bad line hides the clues a person might catch.

When the score crosses the threshold, Sierra flags the call as likely AI. From there, it is your call.

Flow of one phone call through fleming-1: incoming call, live audio scoring, a conservative likely-AI flag, then your rules decide to verify, route or just count it
Flow of one phone call through fleming-1: incoming call, live audio scoring, a conservative likely-AI flag, then your rules decide to verify, route or just count it

Two design choices stand out.

It is tuned conservative. Sierra says the model is set so "real people don't inadvertently get flagged." That is the right default for a support line, because flagging a real customer as a bot is a much worse experience than missing a bot. The trade-off is that some AI callers will pass unflagged. So treat a missing flag as "unknown", not "human".

It reports, it does not act. Sierra leans on the caller ID comparison: "Caller ID never told you whether to pick up." The model gives your Sierra agent a signal, and the agent's own logic, which you write, decides what to do with it.

What Sierra has not said is just as important for anyone planning around it. There are no published figures for accuracy, false positive rate, how many seconds of audio the model needs before it scores, which languages it covers, or which voice generators it was tested against. Compare that with Pindrop, which publishes a 5,000+ AI voice registry and makes accuracy claims on its page. If you are a Sierra customer, ask your account team for numbers on your own recordings before you write policy that depends on the flag.

How fleming-1 fits with the Personal Agent Protocol

fleming-1 shipped the day after Sierra and Meta announced the Personal Agent Protocol, and the two are meant to be read together. Sierra's post puts it plainly: "The best case is an agent that says who it is, and that's what Personal Agent Protocol is for. For the calls where it doesn't, there's fleming-1."

The difference is in what you learn about the caller.

Comparison of what a business learns: with the Personal Agent Protocol it knows the caller is an agent and who it acts for; with fleming-1 it knows the caller is likely an agent but not who it acts for
Comparison of what a business learns: with the Personal Agent Protocol it knows the caller is an agent and who it acts for; with fleming-1 it knows the caller is likely an agent but not who it acts for

The protocol, which I covered in the PAP breakdown, runs on OAuth. A personal agent starts a session, the customer decides whether it gets read or write access, and the business decides what agents may do through its website, its APIs (MCP or OpenAPI) or its own agent (see MCP for customer support). When an agent uses it, "the company knows it's an agent and who it's acting for."

fleming-1 knows much less. It can tell you a voice is probably synthetic. It cannot tell you whose agent it is, whether the customer approved the call, or whether the caller is a person using text-to-speech. That is why Sierra frames it as information. A "likely AI" flag is a reason to slow down on sensitive actions, not proof of fraud.

Sierra is not the only vendor building both halves, and it competes with Parloa and Cresta on voice too. Decagon announced its own Personal Agent Gateway on October 1, which detects personal agents in chat and voice and routes them to a separate channel, plus the PACT protocol for agents to prove who they represent. Decagon has since said it is joining the PAP working group.

Who can use fleming-1, and what it costs

Short version: Sierra customers running voice agents. Sierra's post says "fleming-1 works with any voice agent built on Sierra. You just need to turn it on." There is no standalone fleming-1 API, no listing for other contact center platforms, and no self-serve signup.

Sierra Summit 2026 recap, where Sierra introduced Fleming and Curie as new models in its Agent OS, as taken from Sierra's blog
Sierra Summit 2026 recap, where Sierra introduced Fleming and Curie as new models in its Agent OS, as taken from Sierra's blog

On price, Sierra has said nothing specific to fleming-1. That fits how Sierra sells in general: it publishes no price list, and its contracts are built around outcomes, with consumption-style pricing where outcomes do not fit. I would expect fleming-1 to sit inside your existing Sierra agreement rather than show up as its own line item, but that is an expectation, not something Sierra has confirmed. The Sierra AI pricing guide has what is publicly known about the contracts.

Sierra's reach makes the feature matter even with that limit. The summit recap says Sierra works with almost 50% of the Fortune 50, one in three of the leading banks and 80% of the Fortune 50 healthcare companies. Banks and healthcare are exactly the two places Sierra's post uses as examples, and they are the industries where a synthetic caller asking for an account change is a real risk. If that is your world (or you are weighing Sierra vs Zendesk), the AI customer service for fintech and healthcare guides cover the wider rollout questions.

What to do with a flagged call

The model is the easy part. The hard part is deciding what a flag should change, and Sierra leaves that to you on purpose. Its two examples are a good starting frame: "A bank might want to add a verification step when the caller is an agent. A company with high call volume might start by measuring how often it happens."

I'd treat the options as a ladder and climb it only as far as your data justifies.

Four-step ladder for handling a flagged call: measure it, verify it before account changes, route it to a lane built for agents, and block it only for clear abuse
Four-step ladder for handling a flagged call: measure it, verify it before account changes, route it to a lane built for agents, and block it only for clear abuse
  1. Measure it. Turn the flag on and change nothing for a few weeks, the same way you would baseline any call center automation. Count how many calls are flagged, which intents they carry (balance check, cancellation, refund), and how they end. Most teams have no idea what share of their calls are agents today.
  2. Verify it. Add a step only where the action is sensitive: changing an address, moving money, cancelling a plan. A balance lookup by a flagged caller is low risk; a payout is not.
  3. Route it. If agent calls turn out to be a big, legitimate share, give them their own path. A shorter flow with fewer pleasantries and clearer confirmations suits an AI caller better than a script written for people.
  4. Block it. Keep this for clear abuse, like a burst of flagged calls testing account security. Blocking a whole class of callers will also block customers who use text-to-speech, and customers who sent an agent on their behalf.

The step that does the most work is not about detection at all. It is having your money-moving rules written down in the first place, so a flag tightens an existing policy instead of inventing one under pressure. Clear escalation rules do most of the work. I see this every week in eesel's own rollouts. One digital-media support admin taught their Zendesk agent a rule that applies whether the requester is a person or a bot:

"I have a rule in CS where we do not address a cancel or refund request when there is an issue attached to it." / "This is incorrect. You have not provided troubleshooting steps yet."

A digital-media support admin encoding a "troubleshoot before you cancel" policy into the Zendesk agent

A personal agent asking for a refund should hit that same rule. If your AI agent handoff and refund request policies are clear, a "likely AI" flag becomes one more input, not an emergency.

How fleming-1 compares to other ways to detect AI callers

fleming-1 is one of three main approaches on the market right now. They differ most in who can buy them.

Sierra fleming-1Decagon Personal Agent GatewayPindrop BotStopper
What it isA model inside Sierra's voice agentsA detection layer plus separate channel for personal agentsStandalone AI and bot caller detection
ChannelsPhoneChat and voicePhone
Who can use itSierra customersDecagon customersEnterprises on supported contact center platforms, no other Pindrop product needed
SignalsLive audio scoringDevice fingerprints, account history, conversation patterns, request cadenceLive audio plus a registry of 5,000+ known AI voices
Tells you who the agent acts for?NoThrough the PACT protocol, when agents use itCan match a known registered AI voice
Published accuracyNoneNoneAccuracy and false positive claims on its page
Public priceNoNoNo
AnnouncedOctober 2026October 1, 20262026
Pindrop BotStopper product page, which detects AI and automated callers and checks them against 5,000+ AI voices, as taken from Pindrop
Pindrop BotStopper product page, which detects AI and automated callers and checks them against 5,000+ AI voices, as taken from Pindrop

If you already run Sierra, fleming-1 is the obvious first step. It is built into the agent answering your calls, so there is nothing to integrate, and the flag lands where your call logic already lives.

If you run Decagon, the gateway does a similar job with a broader signal set and covers chat as well. The Decagon vs Sierra comparison covers the platforms more broadly.

If you build on developer platforms like Retell AI or Vapi, none of these three plug in directly, so caller detection is something you add yourself.

If you run neither, Pindrop is the option you can actually buy on its own, and it is built for supported contact center platforms. It is also the most fraud-oriented of the three, with a $1M deepfake warranty available on eligible full-suite contracts.

For the wider set of phone tools, including whether AI can answer support calls at all, the best AI voice agents roundup and the guide to AI phone support go deeper.

The part fleming-1 does not cover: agents in your ticket queue

Phone is only one door. Personal agents use chat and email too, often when the phone line fails. That is why the Instinct and Muse support guides spend as much time on tickets as on calls. One Muse user on Reddit described exactly that: the agent "discovered the humans weren't home, went to chat, and successfully made the correct request" to replace a toll transponder, per u/intenost on Reddit. fleming-1 never sees that conversation, because it only listens to calls.

No helpdesk I know of reliably detects an AI-written email today, and I would be wary of anyone who claims to. The practical defence in text channels is the same as step 2 of the ladder above: decide which actions need a person, and make that rule hold no matter who sent the message.

That is the part eesel is built for. eesel does not do voice, and it does not detect personal agents. Its AI helpdesk teammate joins your existing Zendesk, Freshdesk or Gorgias queue and works tickets like a new hire, tagging and routing with ticket classification the same way your team does.

Every action it can take is set to Auto, Needs approval or Disabled, per the actions and approvals docs. So lookups run on their own, while refunds and cancellations wait for a person to approve, whether the request came from a customer or from their agent.

Try eesel

If Sierra's launch has you thinking about what customers' AI agents will ask your support team to do, start with the queue you already have. eesel's helpdesk teammate plugs into Zendesk, Freshdesk or Gorgias in minutes, learns from your past tickets and help center, and lets you put any money-moving action behind an approval before it goes live. You can test it against your own past tickets first, and the free plan includes 100 credits with no card. Try eesel.

eesel activity view listing Zendesk conversations the AI teammate handled, with pending and resolved status
eesel activity view listing Zendesk conversations the AI teammate handled, with pending and resolved status

Frequently Asked Questions

What is Sierra fleming-1?
Sierra fleming-1 is a model that scores the audio of a live phone call for signs the caller's voice was generated by AI, then flags calls that are likely an AI agent. Sierra announced it at its October 2026 Summit as part of the model set behind Sierra's agent platform. It gives you information, not a verdict: your business decides what happens to a flagged call.
Does fleming-1 block AI callers?
No. Sierra frames it like caller ID: it tells you what is calling and leaves the decision to you. A bank might add a verification step, a high-volume team might only count the calls. Sierra also tuned it to be conservative by default so real people are not flagged by mistake, which means some AI callers will get through unflagged. Plan your escalation rules with that in mind.
How much does Sierra fleming-1 cost?
Sierra has not published a separate price for fleming-1. Its post says the model works with any voice agent built on Sierra and you only need to turn it on. Sierra itself has no public price list and sells mostly on outcome-based contracts, so the real number sits inside your Sierra deal. The Sierra AI pricing breakdown covers what is known.
Can I use fleming-1 without Sierra?
Not today. Sierra describes fleming-1 as working with voice agents built on Sierra, and there is no standalone API or listing for other phone systems. If you run a different contact center platform, standalone detectors such as Pindrop BotStopper are the closer fit, and the Sierra AI alternatives list covers other agent platforms.
How is fleming-1 different from the Personal Agent Protocol?
The Personal Agent Protocol is for agents that announce themselves: the business learns it is an agent and who it acts for. fleming-1 covers the agents that do not follow the protocol. It can say a caller is likely AI, but not whose agent it is or what it is allowed to do.
How accurate is Sierra fleming-1?
Sierra has not published accuracy, false positive rates, latency or supported languages for fleming-1, and I could not find independent testing as of October 11, 2026. The only performance claim is that it is tuned conservative by default. Ask for numbers on your own call recordings before you build policy on top of the flag.
What should a support team do about AI agents contacting them?
Start by measuring how often it happens, then put extra checks only on actions that move money or change accounts. That applies to chat and email too, where agents like Meta Muse also show up. An AI helpdesk teammate with per-action approvals, such as eesel, can keep refunds behind a person regardless of who sent the request.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Blue waveform icon on an abstract blue and purple background
Guides

OpenAI Audio Speech API: what to build and test in 2026

Understand the OpenAI Audio Speech API for text-to-speech, transcription, and realtime voice work, then separate the model infrastructure from the support workflow you operate.

Rama AdiRama AdiOct 12, 2025
Hand-drawn illustration of a customer with an AI-agent earpiece shaking hands with a shopkeeper who holds a ring of keys at an open door
Trending

Personal Agent Protocol (PAP): what Sierra and Meta have shipped so far

Sierra and Meta announced Personal Agent Protocol on October 6, 2026. Here is what is actually published, what is only promised, and how it compares to PACT, A2A and Visa TAP.

Rama AdiRama AdiOct 9, 2026
Hand-drawn hero banner of a person at a laptop sending documents into a small decision box that sorts them into three labelled trays while a second person looks on
Trending

Strands Decider 2B: AWS's free decision model, tested against the hype

Strands Decider 2B is a free, open 1.9B decision model from AWS's Strands Labs. What it does, how accurate and fast it is, what it costs to run, and where it fits.

Rama AdiRama AdiOct 8, 2026
Hand-drawn illustration of people talking with expressive voice bubbles and sound waves, representing the Eleven v4 text to speech model
Trending

Eleven v4: what changed, what it costs, and when to use Turbo

Eleven v4 is ElevenLabs' new text to speech model. Here's what changed from v3, what v4 and v4 Turbo cost, and which one to pick for your project.

KiraKiraOct 6, 2026
Two contrasting AI customer support concierge desks, one modern and one classic, representing a Lorikeet vs Sierra comparison
Trending

Lorikeet vs Sierra: which AI customer support agent wins in 2026?

A hands-on Lorikeet vs Sierra comparison: pricing, how each agent is built, who each fits, and where a lighter helpdesk teammate makes more sense.

Rama AdiRama AdiSep 22, 2026
Illustration of a satellite orbiting the Moon, feeding stacked crater, ice and volcanic data layers into a foundation model that two scientists study
Trending

NASA-IBM Lunar Foundation Model: the open-source Moon AI, explained

NASA and IBM just open-sourced a foundation model trained only on Moon data. It beats a bigger generic model at finding ice and craters. Here is what it is, and the lesson underneath it.

KiraKiraSep 11, 2026
Hand-drawn illustration of two people in a natural back-and-forth voice conversation, in OpenAI teal
Trending

GPT-Live-1 review: is OpenAI's full-duplex voice model worth it?

A hands-on GPT-Live-1 review: OpenAI's full-duplex voice model is the most natural I have used, but it delegates reasoning to a backend model and bills per minute. Here is the honest verdict.

KiraKiraSep 11, 2026
Two people having a natural conversation with an AI voice assistant, sound waves flowing between them
Trending

GPT-Live-1: OpenAI's full-duplex voice model, explained

What GPT-Live-1 actually is: OpenAI's full-duplex voice model that listens and speaks at once, now in ChatGPT and the API at $0.05 per minute.

KiraKiraSep 11, 2026
Two people talking across a table while an audio-visual AI model watches, listens and speaks in the same loop
Trending

SeedRealtime: what ByteDance's audio-visual model actually does

SeedRealtime is ByteDance's audio-visual full-duplex model. Here is what it does, what ByteDance published, and what you can actually call today.

KiraKiraAug 18, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free