Gemini 3.8 Flash TTS review: is Google's new voice model worth it?

Alicia Kirana Utomo
Written by

Alicia Kirana Utomo

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 24, 2026

Expert Verified
Illustration of sound waves, a voice-design control panel and a play button representing Gemini 3.8 Flash TTS

What Gemini 3.8 Flash TTS actually is

First, the boring but load-bearing facts, because half the confusion online is people mixing up model IDs.

Gemini 3.8 Flash TTS is a text-to-speech model: text goes in, audio comes out, nothing else. It shipped as stable, not preview, under the ID gemini-3.8-flash-tts, and Google lists it as the recommended replacement for the older gemini-3.1-flash-tts-preview. There is a cheaper sibling, gemini-3.8-flash-lite-tts, that I will come back to.

A few specs worth pinning down, all from Google's own speech generation docs:

  • 130 languages on Flash TTS, with the input language auto-detected.
  • 8,192 input tokens and 16,384 output tokens per request, so this is built for turns and paragraphs, not whole novels in one call.
  • Default output is 24 kHz mono WAV, with mulaw and alaw options for telephony.
  • Audio bills at 25 tokens per second, which is the detail that makes the pricing math actually work.

It does not do function calling, structured output, thinking, search grounding, or the Live API. This is a rendering engine, not an agent. Keep that framing, because it matters later.

How good does it actually sound?

This is the only question that matters for a voice model, and it is also where the marketing and the reality diverge a little.

Google's case is strong. The launch post claims #1 on Hume AI's Voice Design Benchmark (71.4) and #1 and #2 on Hume's Overall Quality Index for Flash and Flash-Lite. Those are real numbers. One caveat I would want in front of me before I trusted them: Hume's founder is credited as an author on Google's own announcement, so the vendor and the benchmark author are not fully independent here. That does not make the scores wrong, it just means I weight them a bit lighter.

The independent read is more useful. Per Artificial Analysis, Gemini 3.8 Flash TTS took #1 on Pronunciation Robustness at 89.5% (up from 3.1 Flash TTS's 88.2%), but only #2 on the blind Provider Voice Arena, at an Elo around 1,263, behind Sonic 3. So the fair summary is: best-in-class at saying words correctly, second place at sounding like the voice you would actually pick.

"Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. Gemini 3.8 Flash TTS debuts at #1 on our Pronunciation Robustness Benchmark and #2 on our Provider Voice Arena Leaderboard"

Developers who actually ran it were blunter. The sharpest note from the launch-day thread on Hacker News:

Hacker News

"The voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this."

That tracks with the benchmark gap. Pronunciation is solved; personality is not. If your job is a clear, correct, neutral narrator, you will be very happy. If you need a voice that a listener remembers, temper expectations.

What you can actually control

The reason to reach for Gemini over a plain neural voice is control, and this is the part I liked most. You are not just picking a voice off a shelf, you are directing it.

Diagram of Gemini TTS controls feeding into one finished audio scene
Diagram of Gemini TTS controls feeding into one finished audio scene

Four layers stack together:

  1. Voice design. Describe a persona in plain language and the model builds a voice from the description, which you can then save as a reusable ID.
  2. Stage directions. A speech_metadata style field lets you steer tone, accent and pace per turn ("read this warmly, slightly rushed").
  3. Two-speaker dialogue. A conversational mode handles up to two prebuilt voices in one request, which is what makes quick podcast-style clips possible.
  4. Nonverbal tags. Inline markers like <laugh>, <sigh>, <breath> and <short pause> drop vocal events exactly where you want them.

The voice design flow is the headline feature and it is legitimately clever: you go from a sentence of description to a persistent, reusable voice without ever recording anything.

Pipeline showing voice design from a text description to a reusable voice ID
Pipeline showing voice design from a text description to a reusable voice ID

The honest limit: controllability is real but not perfectly reliable. On an earlier Gemini TTS release, one developer reported that direction was sometimes ignored outright:

Hacker News

"No matter what I wrote in the audio profile, AI Studio never followed it, regardless of scene or context. For example, I tried to get a male voice and kept getting female ones."

Google says 3.8 improves creative direction, and in my reading it does, but budget for a few regenerations when a direction does not land the first time. It is a prompt, not a guarantee.

Voice replication (cloning) also shipped, which for a while Google seemed hesitant to do. As Simon Willison put it:

Hacker News

"I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it."

Flash vs Flash-Lite: which one to use

The two-model split is simple once you see it, and picking wrong just costs you money or quality.

Side-by-side comparison of Gemini Flash TTS and Flash-Lite TTS
Side-by-side comparison of Gemini Flash TTS and Flash-Lite TTS

Flash TTS is the expressive one: 130 languages, $9.00 per 1M audio tokens, and the full creative-direction toolkit. Reach for it when the audio is the product, like an audiobook, a branded ad, or a character.

Flash-Lite TTS is the workhorse: 101 languages, $6.00 per 1M audio tokens, tuned for high-volume rendering where "good and cheap" beats "expressive." Reach for it when you are generating thousands of clips and nobody is going to notice a slightly flatter read.

My rule of thumb: prototype on Flash so you hear the ceiling, then drop to Flash-Lite for anything you are rendering in bulk. If neither fits, the Gemini alternatives worth trying are Cartesia's Sonic 3 and ElevenLabs.

The price reality (and the 2027 catch)

I keep the full numbers in the Gemini 3.8 Flash TTS pricing post, but the review-relevant version is short.

At the standard tier, Flash TTS is $9.00 per 1M audio output tokens. Because audio bills at 25 tokens per second, an hour of speech is about 90,000 tokens, so roughly $0.81 per hour of generated audio. Batch and Flex are 50% off that. There is a real free tier through Google AI Studio.

The catch you have to plan around: those are promotional rates through 31 December 2026, and they roughly double on 1 January 2027 (Flash goes to $18.00 per 1M audio tokens). If you are budgeting a 2027 workload, budget the doubled number, not the launch number. This is the same pattern across the Gemini pricing lineup, so it is not a surprise if you follow Gemini closely.

On cost, the community verdict was clear even from skeptics. Simon Willison, testing the conversation mode:

Hacker News

"The conversation mode is neat, and most of my experiments have cost less than a cent."

How it stacks up against the competition

Cross-vendor TTS pricing is a mess because the billing units differ: Amazon, Azure and Deepgram bill per character, while OpenAI and Google bill per audio token. To compare fairly I normalized everything to dollars per hour of audio (using Amazon's own anchor of roughly 23 hours of speech per 1M characters). Treat these as estimates, since speaking rate changes the math.

ModelHeadline price~ $/hour of audioFree tier
Gemini 3.8 Flash TTS$9.00 / 1M audio tokens~$0.81Yes (AI Studio)
Gemini 3.8 Flash-Lite TTS$6.00 / 1M audio tokens~$0.54Yes (AI Studio)
Amazon Polly Neural$16 / 1M chars~$0.701M chars/mo (12 mo)
Azure Neural / HD Flash$15 / 1M chars~$0.650.5M chars/mo
OpenAI gpt-4o-mini-tts$12 / 1M audio tokens~$0.90None (TTS)
Deepgram Aura-2$30 / 1M chars~$1.30$200 signup credit
ElevenLabs Business"as low as 5c/min"~$3.0010k credits/mo

The takeaway: Gemini 3.8 Flash TTS is priced like a commodity neural voice while behaving like an expressive, generative one. It sits right next to Amazon Polly Neural and Azure on cost, undercuts OpenAI and Deepgram, and is a fraction of ElevenLabs per hour. If you care about the price-to-expressiveness ratio, that is the headline of the whole launch.

One nuance to be fair about: ElevenLabs pricing buys you a voice library and a level of character that the cost table does not capture, which is exactly why teams keep paying for it, as the ElevenLabs reviews show.

Should you use it? A quick decision

Rather than another paragraph of "it depends," here is the call I would make by use case.

Where it falls short

To keep this fair, the real limits I would flag before you commit:

  • The voices lack character. The most-upvoted developer critique was that everything sounds "generically Gemini." Great for utility, weaker for brand.
  • Direction can be ignored. Creative control is a prompt, not a contract, so expect occasional misses and regenerations.
  • It is not the top of the arena. #1 on pronunciation, but #2 on blind human preference behind Sonic 3. If "best voice" is your requirement, it is not settled.
  • Cost doubles in 2027. The launch price is promotional. A workload you size in 2026 gets 2x more expensive in January.
  • Local models are catching up. Several developers pointed to open options like Fish Audio, Higgs and Qwen3 TTS as good-enough for a final render, especially where cost or privacy rules out a cloud API.

None of these are disqualifying. They are the difference between "amazing" and "very good and very cheap," which is where I land.

The part people actually came here for: voice for support

A lot of people searching for a voice model are really asking a different question: can I put this in front of my customers? This is the section I would not skip.

A TTS model is infrastructure. It turns text into audio, and that is the whole job. It does not know your refund policy, cannot read your customer service software, and has no idea whether the sentence it just voiced is actually true. If you point a raw voice model at a support line, it will confidently read out a wrong answer in a lovely voice, which is worse than a plain error message. That is the opposite of the cost savings most teams are chasing.

I work on this at eesel, and the way I think about it is simple: the model is the voice, the teammate is the employee. eesel is an AI teammate platform, and the relevant teammate here is the AI helpdesk agent. It plugs into your existing helpdesk, learns from your past tickets and help center, follows the rules you set, and gets simulated against your real ticket history before it ever answers a live customer. That last part matters: we have watched confident-sounding bots quietly give wrong answers, which is exactly why every rollout gets dry-run on historical tickets first. A voice model has no equivalent safety net.

If you work in a terminal or want to script this, the eesel CLI drives the same teammate and workspace as the dashboard. A person can inspect a teammate's instructions from the command line, a script can automate a workflow, and a coding agent like Claude Code, Codex or Cursor can operate eesel programmatically, print JSON results, and propose bounded changes for an owner to approve. It is the agent-friendly surface for the same product, not a separate thing. That is a real difference from wiring up Gemini TTS yourself, where you own the entire orchestration, the support metrics, the escalation logic and the SLA handling on your own.

So: use Gemini 3.8 Flash TTS for what it is great at, which is cheap, clear, controllable speech. For a voice that actually resolves customer problems, you want a tested teammate on top of it, not the raw model. If that is the real goal, start with a proper AI helpdesk chatbot and the broader support automation tooling, and compare the best AI agent for customer service options before you build anything from scratch.

My verdict

Gemini 3.8 Flash TTS is a very good voice model at a genuinely disruptive price. It nails pronunciation, gives you real creative control, ships voice design and cloning, and costs a fraction of ElevenLabs per hour. The honest asterisks are that the voices lack standout character, it is #2 rather than #1 on blind preference, and the price doubles in 2027.

For most developers rendering audio at scale, it is the new sensible default. For a signature brand voice, keep ElevenLabs and Sonic 3 in the running. And if you came here wanting AI that talks to your customers and actually solves their problems, remember that the voice is the easy part: the AI customer service companies worth your time are the ones solving the hard part, and the AI helpdesk software that resolves a ticket matters far more than the voice that reads the answer.

eesel AI helpdesk dashboard resolving customer tickets
eesel AI helpdesk dashboard resolving customer tickets

Want an AI teammate for your support queue, not just a nice voice? eesel plugs into your helpdesk in minutes, learns from your own tickets, and you can simulate it on your real history before it goes live. Try eesel free, or book a demo to see it on your own data.

Frequently Asked Questions

Is Gemini 3.8 Flash TTS good?

Yes, with a caveat. It is studio-grade on clarity and pronunciation, and it debuted at #1 on Artificial Analysis's Pronunciation Robustness benchmark. But on blind human preference it landed #2, behind Cartesia's Sonic 3, and several developers said the prebuilt voices sound generic. It is an excellent default, not a clean sweep.

How much does Gemini 3.8 Flash TTS cost?

It is $0.50 per 1M text input tokens and $9.00 per 1M audio output tokens on the standard tier through 31 December 2026, which works out to roughly $0.81 for an hour of speech. Both rates double on 1 January 2027. Full tiers are in my Gemini 3.8 Flash TTS pricing breakdown, and the wider family sits in the Gemini 3 pricing guide.

Gemini 3.8 Flash TTS vs ElevenLabs: which is better?

Gemini wins on price by a wide margin, roughly 4 to 5 times cheaper per hour of audio. ElevenLabs still tends to win on standout, characterful voices and its voice library. If cost and scale matter most, Gemini; if a distinctive voice matters most, look at ElevenLabs alternatives too.

What is the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?

Flash TTS covers 130 languages at $9.00 per 1M audio tokens and is built for expressive, directed audio. Flash-Lite TTS covers 101 languages at $6.00 per 1M audio tokens and is built for high-volume, cost-sensitive jobs. Pick Flash when creative direction matters, Flash-Lite when you are rendering at scale.

Can Gemini 3.8 Flash TTS clone a voice?

Yes. It supports voice design (describe a persona in words) and voice replication (clone from a short reference clip plus consent audio). Cloned voices can be stored as a reusable voice ID, capped at 200 per project with a one-year TTL, or as a 7-day stateless key.

Does Gemini 3.8 Flash TTS support multiple speakers?

Up to two speakers per request using prebuilt voices, with a conversational mode for dialogue. For more than two voices you synthesize each turn separately and stitch the audio together.

Is a voice model like Gemini TTS enough for customer support?

No. A TTS model turns text into speech; it does not read your help center, follow your policies, or resolve a ticket. For support you need an AI teammate layered on top, like an AI helpdesk agent that connects to your customer service software and is tested on your real ticket history first.

Share this article

Alicia Kirana Utomo

Article by

Alicia Kirana Utomo

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Illustration of sound waves and price tags representing Gemini 3.8 Flash TTS pricing
Trending

Gemini 3.8 Flash TTS pricing: what an hour of AI voice actually costs

Gemini 3.8 Flash TTS is $9 per 1M audio tokens, roughly $0.81 an hour of speech. Here is every tier, the 2027 catch, and how it stacks up against ElevenLabs and OpenAI.

Rama Adi NugrahaRama Adi NugrahaSep 24, 2026
Illustration of sound waves and price tags representing Gemini 3.8 Flash TTS pricing
Trending

Gemini 3.8 Flash TTS pricing: what an hour of AI voice actually costs

Gemini 3.8 Flash TTS is $9 per 1M audio tokens, roughly $0.81 an hour of speech. Here is every tier, the 2027 catch, and how it stacks up against ElevenLabs and OpenAI.

Rama Adi NugrahaRama Adi NugrahaSep 24, 2026
Two people talking across a table while an audio-visual AI model watches, listens and speaks in the same loop
Trending

SeedRealtime: what ByteDance's audio-visual model actually does

SeedRealtime is ByteDance's audio-visual full-duplex model. Here is what it does, what ByteDance published, and what you can actually call today.

Alicia Kirana UtomoAlicia Kirana UtomoAug 18, 2026
Illustration of a team reviewing Gemini 3.8 Flash, with a speed gauge, a rocket, and a verdict checkmark
Trending

Gemini 3.8 Flash review: fast, verbose, and not the upgrade the number implies

A hands-on Gemini 3.8 Flash review: what it's good at, where it falls down, the 13-second catch nobody quoted, and whether to switch from 3.7 Flash.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 8, 2026
Illustration of two people reviewing charts and speed dials around a Gemini spark, representing Gemini 3.8 Flash pricing
Trending

Gemini 3.8 Flash pricing: every rate, the hidden cost, and the catch

Gemini 3.8 Flash costs $0.75/$3.75 per 1M tokens, exactly what 3.7 Flash costs. But the sticker price hides a verbosity tax, and both numbers double on 1 January 2027.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 8, 2026
Illustration of a fast-moving robot coding on a laptop while a person watches, representing Gemini 3.8 Flash
Trending

Gemini 3.8 Flash: what it is, honest benchmarks, and my review

Google shipped Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7. Same price, better scores, and one line of fine print that changes the answer.

Alicia Kirana UtomoAlicia Kirana UtomoSep 3, 2026
Gemini 3.5 Pro review hero banner in Google blue
Trending

Gemini 3.5 Pro review: the honest state of Google's flagship

An honest Gemini 3.5 Pro review: it isn't out yet. Here's what Google has confirmed, why it's late, the benchmarks that do exist, and what to use today.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieJul 21, 2026
GPT-Live review hero illustration, OpenAI's real-time full-duplex voice AI for ChatGPT
Trending

GPT-Live review: is OpenAI's new voice AI worth it?

A hands-on review of GPT-Live, OpenAI's new full-duplex voice model for ChatGPT: what's good, what's missing, and whether it's worth it for support teams.

Riellvriany IndriawanRiellvriany IndriawanJul 13, 2026
GPT-5.6 versus Gemini 3 comparison hero illustration, two AI model families balanced against each other
Trending

GPT-5.6 vs Gemini 3: which AI model wins in 2026?

GPT-5.6 vs Gemini 3 compared: Sol, Terra and Luna against Gemini 3.5 Flash and 3.1 Pro on pricing, benchmarks, context, and which fits AI support agents.

Rama Adi NugrahaRama Adi NugrahaJul 10, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free