Conversational AI for customer service: what actually works in 2026

Riellvriany Indriawan
Written by

Riellvriany Indriawan

Katelin Teen
Reviewed by

Katelin Teen

Last edited July 27, 2026

Expert Verified
Illustration of a support agent working alongside an AI assistant handing a conversation to a customer

What conversational AI for customer service actually means now

A chat widget with a decision tree behind it, that was the old definition. You built the branches, the customer picked one, and if their question sat off the tree they got a "let me connect you to an agent" and a queue. Plenty of those are still running. It is why so many people flinch at the phrase in the first place. What shipped since is different enough that I keep a separate explainer on AI agents versus chatbots.

What ships in 2026 differs in three specific ways.

It retrieves instead of matching. The system reads your help centre and your past tickets, your internal docs too, then composes an answer for the question in front of it. Zendesk runs this across 80 languages on a multi-model stack, and the same retrieval approach underpins every serious AI customer service chatbot on the market.

It takes actions. A modern agent can look up an order or start a return, and apply a credit through an API call too. That is the jump from answering to resolving, which is the point where AI for refund requests stops being a demo. The same shift shows up in order-tracking automation.

It works inside the helpdesk you already have. The bot no longer needs its own inbox. It writes internal notes, drafts replies, tags, and routes, which is how AI ticket classification turned from a side feature into table stakes. Ticket sentiment analysis went the same way.

Decision-tree chatbotConversational AI agent (2026)
Answer sourceHand-built branchesHelp centre, past tickets, internal docs
Off-script questionDead end or handoffComposed answer, or an honest "I don't know"
ActionsUsually noneOrder lookups, refunds, account changes via API
Where it livesIts own widgetInside Zendesk, Freshdesk, Gorgias, Slack, email
Setup workWeeks of branch buildingConnect knowledge, then scope what it may touch
Failure modeFrustrating loopsConfident wrong answers

That last row is the one to sit with. The failure mode moved. A tree bot fails loudly and the customer knows it, but a conversational AI agent fails quietly and fluently, which is exactly why the control layer further down matters more than the model. If you want the wider category view before the money talk, the roundup of conversational AI platforms covers who is building what.

The part vendors leave off the pricing page

Six major vendors, six different billing units, and none of them interchangeable. This is the thing I wish someone had put in front of me two years ago.

Infographic showing one customer conversation fanning out into five different billing meters: per agent seat, per automated resolution, per AI session, per credit and per ticket
Infographic showing one customer conversation fanning out into five different billing meters: per agent seat, per automated resolution, per AI session, per credit and per ticket

The published numbers, all pulled from vendor pricing pages this week:

VendorAI billing unitPublished AI rateWhat the unit countsSeat cost on top
eeselTicket or chat session$0.40One ticket or one chat session, no matter how many repliesNone
Help ScoutResolution$0.75Conversation resolved with no escalation, one per conversation$25 to $75 per user/mo
GorgiasAutomated interaction$1.50Any AI interaction past the plan allowance (30 to 530)Plan fee $40 to $1,430/mo
FreshdeskFreddy AI Agent session$0.49Billed per session after the first 500, sold as $49 per 100$19 to $89 per agent/mo
AgentforceConversation, or Flex Credit$2.00, or $0.005/creditOne conversation, or roughly 20 credits per action$5 per user/mo, plus Service Cloud
ZendeskAutomated resolutionNot publishedLLM-verified resolution after 72 hours of silence$55 to $115 per agent/mo

Four things in that table are worth more than the rest of this section.

Gorgias bills interactions, not resolutions. At $1.50 the headline looks mid-pack, though an "automated interaction" turns out to be a broader unit than a resolution, which means a conversation the AI touched and never solved can still meter. Help Scout goes the other way and only counts a resolution when the customer neither escalates nor clicks "I still need help". Same-ish price, different thing being sold.

Zendesk documents everything except the number. Its help centre spells out the automated resolution mechanic in detail: 5 to 15 resolutions bundled per agent per month, a hard cap of 10,000 allocated resolutions a year on every plan. Overage gets priced explicitly above committed usage, on top of that. The dollar figure appears nowhere. Every path ends at a sales call, which is the same wall readers hit in my Zendesk AI writeup.

Salesforce's own worked examples do not reconcile. The Agentforce pricing page publishes five scenarios at $0.005 per credit, and four of them are internally consistent. The Voice row prints 120 credits at $0.15 when the same maths gives $0.60, because Voice runs on different multipliers. Salesforce says so in its FAQ, and points to a separate rate card. Fine, though it means the published examples are not really a quote, a point my Agentforce pricing guide makes at more length.

The outcome-pricing leaders publish no price at all. Sierra markets outcome-based pricing under the heading "Pay for a job well done" with no unit and no rate anywhere on the site, and Decagon does not state a billing model on its homepage either. Both are real products with real deployments, which my Decagon review goes into at length. There is a matching writeup on Sierra too. They are just not shoppable without a call.

Run your own numbers

The spread only becomes obvious once your actual volume is in it, which is why there's a calculator below, built off the published rates above.

Not a rounding error, the spread between cheapest and priciest at a thousand conversations a month. It is the cost of another headcount, and entirely a function of which noun the vendor decided to meter.

Buyers tend to notice this on the way up, not on the way in:

G2

"The only thing I'd keep an eye on is the pricing. Once you start using more advanced automations or Lyro AI, costs can climb pretty quickly depending on how many conversations you're handling. It's not a deal-breaker, but it's definitely worth understanding how the limits work before you start scaling everything."

What the numbers look like once it is actually live

50% to 90%. That's the resolution-rate range vendor case studies quote. Decagon publishes eight named-brand figures: Duolingo at 80% deflection, Chime at 70% resolution. Zendesk's page carries Best Egg at 80%, Fortnum & Mason at 90%. Real numbers, for those companies. They are also the tail end of a tuned deployment. Month one looks nothing like that.

Here is what a cross-validated trial looked like on a live Zendesk queue at a DTC accessories brand, 284 AI chats checked against 100 real tickets:

Bar chart from a real support trial showing triage accuracy at 93 percent, drafts directionally right at 88 percent, drafts sent with no edit at 12 percent and a factual error rate of 7 percent, with the gap between 88 and 12 labelled the review gap
Bar chart from a real support trial showing triage accuracy at 93 percent, drafts directionally right at 88 percent, drafts sent with no edit at 12 percent and a factual error rate of 7 percent, with the gap between 88 and 12 labelled the review gap

Triage was excellent. Spam detection ran at 100% with zero false positives on an inbox that was 22% spam. Draft quality by category was strong where the category was narrow: returns and refunds 93.8% useful, warranty claims 96.4% useful. Product enquiries hit 100%.

And still only 12% of drafts went out untouched.

That gap between "directionally right" and "send it" is the single most under-discussed number in this category. Not a failure, this. It's the accurate picture of what an unscoped agent does on a mixed queue, and it is why the honest first-month target is a slice of tier-1 rather than a percentage of everything. An internal IT desk on Jira Service Management that I have watched closely sits at 15% deflection against a 55% target, and it is getting there by widening scope one intent at a time, not by swapping models.

The counterweight, and it is a real one: when scope is tight, volume stops mattering. One eesel customer runs a fully automated Zendesk agent entirely in German at over 100,000 tickets a month, per eesel's own customer numbers. Another, a gig-economy driver-analytics app on Zendesk, resolved 73% of its tier-1 requests in the first month after a seven-day trial. Narrow-and-deep, both of them, not broad-and-shallow. If you want the metric definitions before you set a target, start with AI resolution rate, then pair it against AI CSAT so a fast wrong answer cannot look like a win.

One more thing on numbers, because it trips up every comparison I see: deflection and resolution are not the same metric. Decagon reports both in the same ROI band. Deflection means the customer stopped asking, and that includes the ones who gave up. Resolution means they got the answer. My breakdown of ticket deflection keeps the two apart, and so does the piece on resolution-rate reporting. Any vendor table that does not is doing you a small disservice.

What customers say about talking to these things

The other side of the widget, worth reading before you buy one. The public record in 2026 isn't some blanket "AI bad." Much more specific than that.

The case for automating at all is stronger than most vendors bother to make. From someone with two relatives inside support orgs:

Hacker News

"My brother used to work at tech support for XBox Live.

He said that 80% of his calls were for password resets, something users can easily self-service. There's literally an option on the login form for "Forgot Password", and people would rather spend time calling up support, waiting on hold, and verifying their identity to a support agent than click a button. […] I have an uncle that works tech support for XFinity. Half his calls are resolved by just power cycling the modem/router."

That is the volume conversational AI exists to absorb, and it maps almost exactly onto what an AI help desk is good at.

The case against is never really about answer quality. It is about what happens when the answer is wrong and there is no way out:

Hacker News

"The marketing team has 4 turns with an AI that refuses to escalate to a human and is convinced this is the only entry it needs us to add to our DNS. Dropped the vendor. Someone from their retention team followed up and we linked them the ticket talking to the bot about the obvious bug. Never heard back."

Four turns and a cancelled contract. The bot did not lose that deal by being wrong, it lost it by having no exit. Buyer reviews name the same boundary from inside the tooling, with one Tidio reviewer on G2 noting the bot "may require a human takeover sooner than expected" on multi-step questions. That is the honest shape of the technology, and designing around it is the job.

The real blocker is control, not accuracy

Same place, every time. Every stalled deal I have watched stalls right there, and it is never "the AI is not smart enough." It is a support lead working out that they cannot supervise what they cannot scope.

The clearest version of it came from a CX lead at a DTC supplements brand running Gorgias and Shopify, around 7,000 tickets a month:

"The AI will never be able to answer 100% of the questions, but if it tries and just answers 'sorry I don't know this,' I cannot go and check all my 7,000 tickets to see if the AI actually made a good answer. I need an AI who is only handling the tickets that it's confident to handle and all the other ones, leave them alone."

That is the whole objection in one paragraph, and it is a good one. The same theme shows up over and over in the notes: "there are certain tickets I don't want to go through AI", and "I want response only when I mention @eesel, not during creation and every customer ticket message". These are not people who distrust AI. They are people who want a dial. Someone on Hacker News put the same point more bluntly than I ever would:

Hacker News

"The problem is not chatbot customer support, the problem is bird-brained managers that think a system that solves 99% of issues doesn't need a fallback for that 1%."

Infographic showing incoming conversations passing through an excluded-topics gate, then a confidence threshold that splits them into auto-reply and close, or draft only for a human to send
Infographic showing incoming conversations passing through an excluded-topics gate, then a confidence threshold that splits them into auto-reply and close, or draft only for a human to send

The pattern that works is two gates in series. First an exclusion list: billing disputes, legal questions, refunds over a threshold, none of it reaches the AI at all. Then a confidence threshold, so anything the agent is unsure of becomes a draft in the agent's queue rather than a reply in the customer's inbox. Zendesk exposes its own version of this as an intent confidence threshold. The handoff design around it matters as much as the threshold value.

Set it up that way and the 12%-sent-untouched figure stops being alarming, because the 88% that needed edits never went out unsupervised in the first place. Set it up without gates and that same 7% factual error rate lands in front of customers. My writeup on preventing hallucinations goes deeper on the mechanics.

How I would roll one out

Four steps, in this order. The order is the part people skip.

1. Simulate on historical tickets before anything goes live. Run the agent against a few hundred of your own closed tickets and read what it would have said. This is the step that catches the confidently wrong answer while it is still free, and it is the reason eesel builds simulation into onboarding rather than selling it as an add-on. It also gives you a defensible resolution-rate estimate instead of a vendor's. Getting the inputs right first helps, which is what training on your knowledge base is for.

2. Scope narrow, then widen. Pick the two or three intents with the highest volume and the lowest blast radius. Order status and password resets. Shipping windows, maybe. Nothing that touches money or law on day one. Widen only when the numbers on the current slice hold for a fortnight. My guide to deflecting FAQs with AI is a decent starting shortlist.

3. Set the confidence threshold before go-live, not after the first bad reply. Start conservative. A high threshold on a narrow scope produces a boring agent that nobody complains about, which is exactly what you want in week one. You can always loosen it, and the handoff best practices worth copying are mostly about what happens at the moment you do.

4. Watch the unit, not the invoice. Whatever you sign, instrument the thing your vendor meters. If it is resolutions, track the escalation rate that kills them. If it is interactions, track how many the AI touched without solving, since those bill too. Zendesk confirms a resolution only after 72 hours of customer silence plus an LLM check, so its usage dashboard trails reality by three days, which matters if you are managing to a monthly cap.

Where I would not put conversational AI yet

The difference between a rollout that survives and one that gets switched off in month two? Being straight about this.

Anything where a wrong answer costs real money or creates legal exposure should stay human, at least until you have months of data. Regulated advice, disputes, cancellations with retention offers, anything involving a chargeback. The technology can hold the conversation; the point is that the cost of the 7% is asymmetric.

Queues with almost no repeatable volume are also a poor fit. If every ticket is bespoke, there is nothing for retrieval to retrieve, and you will spend more on configuration than you save. Complex escalation paths are worth automating; problems nobody has seen before are not.

And a limit on my own side, since it is only fair: eesel is built to layer onto the helpdesk you already run, which means it is a poor pick if what you actually want is a full replacement helpdesk with ticketing, telephony, workforce management, the whole stack in one bill. That is a real thing to want, and a full-stack helpdesk like Zendesk will serve you better for it. The same goes for Freshworks AI. eesel also has no native Crisp or LiveAgent integration, so on those stacks it is an alternative rather than an add-on.

Try eesel for conversational customer service

That specific thing, the conversational layer on top of a helpdesk you already run, is what eesel does. It plugs straight into Zendesk, reads your existing macros and help centre, your closed tickets too. Freshdesk, Gorgias, Front and Slack work the same way, and most teams have a first agent live inside 30 minutes.

The eesel dashboard showing an AI agent handling live Zendesk ticket activity
The eesel dashboard showing an AI agent handling live Zendesk ticket activity

Two things I would point at specifically, given everything above. You can simulate the agent against your own closed tickets before it replies to a single customer, so the resolution-rate number you take to your manager is yours and not a case study's. And the meter is $0.40 per ticket or chat session on eesel's pricing, with no seat fees and no platform fee, billed per conversation handled rather than per reply, which means a chatty customer does not cost more than a terse one. There is $50 of free usage to test that with, no card required.

If you would rather see the alternatives first, my Zendesk AI alternatives roundup is the honest place to start. The wider customer service AI comparison covers the rest of the field.

Frequently Asked Questions

What is conversational AI for customer service?
It is software that holds a real back-and-forth with a customer, pulls the answer from your help centre and past tickets, and where it is allowed to, takes an action like issuing a refund or checking an order. That last part is what separates it from an old decision-tree bot, and it is the line drawn in AI agents versus chatbots. Most teams meet it first as a customer service chatbot on the website, then extend it into email and chat.
How much does conversational AI for customer service cost?
Published rates in 2026 run from $0.40 per ticket with eesel to $2.00 per conversation with Agentforce, with Help Scout at $0.75 per resolution and Gorgias at $1.50 per automated interaction. Zendesk documents the unit but not the price, which is why my Zendesk pricing breakdown ends at a sales call. Seat fees usually stack on top of all of that.
Is conversational AI in customer service better than a live chat widget?
They do different jobs. A widget routes a human to the conversation, and conversational AI answers it. In practice most teams run both, with the AI on tier-1 and a clear handoff to a person, which is the pattern in multilingual live chat too.
What percentage of tickets can conversational AI actually resolve?
Vendor case studies quote 50% to 90%, but the honest starting range for a first month is closer to 15% to 30% while scope is tight. One internal IT desk on Jira sat at 15% deflection against a 55% target. Read AI resolution rate and how to improve it before you commit to a number in a board deck.
How do I stop conversational AI from giving customers wrong answers?
Three controls do most of the work: exclude the topics it must never touch, set a confidence threshold so unsure conversations become drafts instead of replies, and simulate against historical tickets before it goes live. My guide to AI hallucinations in support covers the failure modes, and training on your knowledge base covers the input side.
Can conversational AI work with my existing helpdesk?
Yes, and layering onto the helpdesk you already run is usually cheaper than migrating to a new one. eesel connects to Zendesk, Freshdesk, Gorgias and others, and the same logic applies to picking AI for Help Scout rather than replacing it.
What is the difference between deflection and resolution in AI customer service?
Deflection means the customer went away, and resolution means the customer got an answer. Vendors mix the two in the same ROI table, which is why deflection numbers and resolution-rate metrics are not comparable. Pair either with AI CSAT before you believe it.
Is there a free way to try conversational AI for customer service?
There is. eesel gives you $50 of usage with no credit card, Help Scout runs three months of unlimited AI Answers resolutions, and Freshdesk includes the first 500 Freddy sessions on every plan. There is a wider list in my roundup of free AI for customer service tools.

Share this article

Riellvriany Indriawan

Article by

Riellvriany Indriawan

Riell is a designer and writer at eesel AI with about two years of experience researching CX platforms, AI chatbots, and helpdesk software. She combines her design background with a sharp eye for how these tools actually look and feel in practice — making her comparisons unusually visual and user-focused.

Related Posts

All posts →
Illustration of an AI chatbot answering student questions across schools and universities
Guides

AI chatbots for education: a practical guide for student support

AI chatbots for education answer student questions 24/7, cut summer melt, and free up staff. Here's what to automate, what to escalate, and how to stay safe.

Riellvriany IndriawanRiellvriany IndriawanJul 15, 2026
Illustration of an AI chatbot and a support agent handling customer conversations in Kustomer
Guides

AI chatbot for Kustomer: how it works, what it costs (2026)

How an AI chatbot works inside Kustomer, what Concierge and Envoy actually do, what you'll really pay, and when a drop-in AI layer beats the native option.

Rama Adi NugrahaRama Adi NugrahaJul 14, 2026
Illustration of an AI chatbot for fintech customer support with a secure chat and finance motif
Guides

AI chatbot for fintech: what works, what breaks in 2026

What an AI chatbot for fintech actually does, why a wrong answer costs more than an annoyed customer, and how to deploy one that survives a security review.

Rama Adi NugrahaRama Adi NugrahaJul 12, 2026
Twig AI: A complete 2025 overview for support teams
Guides

Twig AI: A complete 2025 overview for support teams

ig AI enhances customer support by delivering AI-driven conversations that are natural, efficient, and scalable.

Stevia PutriStevia PutriSep 9, 2025
A provider's guide to conversational AI for healthcare in 2025
Guides

A provider's guide to conversational AI for healthcare in 2025

Discover how conversational AI in healthcare improves patient engagement and streamlines operations, with key use cases and tips for choosing the right platform.

Stevia PutriStevia PutriAug 4, 2025
Serval AI pricing in 2026: How the pilot model works
Guides

Serval AI pricing in 2026: How the pilot model works

Serval doesn't publish per-seat or per-ticket pricing. Here's how the pilot model actually works, sourced from Serval's own pricing page, plus how it compares to public per-interaction pricing.

Alicia Kirana UtomoAlicia Kirana UtomoMay 2, 2026
A practical guide to the best AI tools for IT support in 2026
Guides

A practical guide to the best AI tools for IT support in 2026

Struggling with slow, costly IT support? Explore the top AI tools for IT support and learn how to automate tasks, reduce ticket backlogs, and improve team efficiency.

Stevia PutriStevia PutriNov 13, 2025
A complete guide to Voiceflow pricing in 2025
Guides

Voiceflow pricing 2026: Credits, plans, and is it worth it?

Voiceflow pricing offers flexibility for solo creators and enterprises alike. See how plans scale for building and managing AI-driven conversations.

Stevia PutriStevia PutriAug 25, 2025
Illustrated hero banner for a guide on AI chatbots for the travel industry
Guides

AI chatbot for travel: a practical 2026 guide

What an AI chatbot for travel really does across the trip, what it costs you when it's badly built, and how to pick one that can actually rebook a stranded traveler.

Rama Adi NugrahaRama Adi NugrahaJul 16, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free