Gemini 3.8 Flash review: fast, verbose, and not the upgrade the number implies

Kurnia Kharisma Agung Samiadjie
Written by

Kurnia Kharisma Agung Samiadjie

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 8, 2026

Expert Verified
Illustration of a team reviewing Gemini 3.8 Flash, with a speed gauge, a rocket, and a verdict checkmark

What Gemini 3.8 Flash actually is

Gemini 3.8 Flash is Google's newest fast, cheap, multimodal model, available as gemini-3.8-flash in the Gemini API and as the default model in Google's Antigravity agentic dev environment. It takes text, images, audio, video, and PDFs in, and writes text out, with a 1M-token input window and 64K output. If you want the full launch breakdown, I covered it in my Gemini 3.8 Flash overview, and the money side lives in the Gemini 3.8 Flash pricing post. This one is the review: is it worth your time, and where does it actually fit.

The Gemini 3.8 Flash model page, as taken from Google DeepMind
The Gemini 3.8 Flash model page, as taken from Google DeepMind

The single most useful fact for a review is one nobody put in a headline: this is not a fresh model. The model card states, verbatim and in four separate sections, that "Gemini 3.8 Flash is based on Gemini 3.7 Flash." Architecture, training data, hardware, and safety policy all defer to the 3.7 Flash card. So 3.8 is a post-training pass on 3.7's base. That explains the three-week cadence, and it explains why the price did not move a cent. Read this review as "is the extra post-training worth switching for," because that is the real question.

The one number that decides this review

Every launch post led with speed. Here is the number that reframes it. Artificial Analysis clocks Gemini 3.8 Flash at 305.1 tokens per second, ranked #3 of 195 models. That is real and it is impressive. But the same measurement page also reports a 13.30-second time to first token, where the median reasoning model in its price tier sits at 2.99s.

Gemini 3.8 Flash is third fastest by output speed but takes 13.30 seconds to produce its first token
Gemini 3.8 Flash is third fastest by output speed but takes 13.30 seconds to produce its first token

Those two numbers describe completely different experiences. Throughput is how fast the essay pours out once it starts. Time to first token is how long the reader stares at a blank box before anything appears. For an overnight batch agent chewing through a repo, 13 seconds of upfront thinking is invisible and the throughput is everything. For a person watching a chat widget, or a support agent waiting on a suggested reply, 13 seconds of dead air is the whole experience, and no amount of downstream speed saves it.

This is why I keep pushing back on "3.8 Flash is the fast one". It is the fast one for machines and the slow one for humans. That single distinction sorts almost every use case I can think of, and it is the reason the benchmarks below matter less than they look.

Where it's genuinely good

Give it a long, agentic coding job and 3.8 Flash earns its keep. On Google's own numbers it tops HLE-Verified at 54.9% (ahead of GPT-5.6 Sol, Claude Opus 5, and its own predecessor), wins the Vals Finance Agent v2 benchmark at 61.4%, and sits top of DeepSWE for long-horizon software work. Benchmark-vendor david put the appeal cleanly: "5x cheaper and 2x faster than Opus 5 for similar intelligence."

Where Gemini 3.8 Flash leads on benchmarks and where it trails
Where Gemini 3.8 Flash leads on benchmarks and where it trails

The Intelligence Index backs the coding story up: 59 on Artificial Analysis, matching Claude Opus 5 on medium effort, which for a Flash-class model is a strong result. And the long-horizon behaviour is real, not a benchmark artifact. Design partner Thai Tran of Glean reported it "completing more than three times as many tasks as Gemini 3.7 Flash in our evaluations" on document-heavy workflows. On Hacker News, simonw generated an HTML visualization in 13 seconds for 1.8 cents and called the speed-plus-frontend-quality combination "pretty exciting."

My favourite real-world pattern came from panarky, who uses Flash as an auditor rather than a builder:

Hacker News

"I've been using 3.7 Flash to audit the work of Opus High, and Flash finds lots of subtle and insidious defects even while all the unit tests are green... it is blazing fast in Antigravity CLI. Easily 10x faster than Opus. Can't wait to try 3.8 Flash. If it's good enough, maybe I'll switch Flash to primary and make Opus the auditor."

That is the shape of workload 3.8 Flash is built for: cheap, fast, tireless, willing to grind a long task. If that is you, it belongs on your shortlist next to the other best AI coding assistant tools.

Where it falls down

The same "works harder" design that wins the coding benchmarks is what makes it awkward everywhere else. Google is unusually honest about this in the launch post: 3.8 Flash "exhibits greater diligence, executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance." Translated: it thinks a lot, and you pay for the thinking.

Gemini 3.8 Flash used 120M output tokens against a 71M median to run the same benchmark suite
Gemini 3.8 Flash used 120M output tokens against a 71M median to run the same benchmark suite

Artificial Analysis flagged it as "very verbose": 120M output tokens to run the Index against a 71M median. Since Gemini bills thinking tokens at the $3.75 output rate, that verbosity is a direct cost multiplier the sticker price hides. HN user bermudi put a number on it, roughly 11k more tokens per task than 3.7 on the same benchmark, "making it more expensive and slower in practice than the headline rate suggests." My Gemini 3.8 Flash pricing post has the full arithmetic, but the short version is that a per-1M rate lies when the model chooses to emit more 1Ms.

Then there is the babysitting. mythz captured a frustration I share:

Hacker News

"Whilst it's a fast model, having to baby sit through and approve prompts every few seconds ends up making it slower than the Auto approve modes of Claude/ChatGPT."

And it can overreach on simple prompts. refulgentis reported saying "Hi" and getting a four-panel app with a fake weather widget and a to-do list. It is a small thing, but it is the tell of a model tuned to do maximum work, which is exactly what you do not want on a narrow, well-scoped task.

Two more honest cautions. On Terminal-Bench 4.0 it scores 19.1% against Opus 5's 51.8%, so "matches Opus 5" is true only on the benchmarks Google chose to show. And the model card records two safety metrics that regressed: multilingual safety by 5.4pp and unjustified refusals by 1.1pp, with Google stating plainly that "safety performance across non-English languages regressed slightly." The full Frontier Safety assessment was not even re-run; it was inherited from 3.7 by inference. None of these are dealbreakers on their own, but they are the kind of detail a fair review has to name.

What people actually said after running it

The community read is split in a useful way, and it maps almost exactly onto the machine-versus-human divide above.

The people running long coding jobs mostly liked it. Selta called it "quite a big update," noting the coding-reasoning gains and reduced hallucinations while keeping the speed. Conor Dart said "Google is back... this model takes its time to think and work through the tasks." And scrlk raised a point benchmarks miss entirely: "Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic." Consistency under load is a real advantage.

The people expecting a leap were cooler. markasoftware noted that on Artificial Analysis it only matches "opus 5 medium effort," and that "opus 5 medium outputs 4x fewer tokens to achieve the same result, negating a lot of the speed difference." zuzululu landed where I did: "i might consider 3.8 flash for simple side hobby projects or quick scaffolding but would not trust it for long agentic tasks." And pkos98 offered the healthiest skepticism about any launch-day benchmark: "Wait a week with your judgement, most likely, Google is just bench-maxing very hard."

There is also a structural gripe worth surfacing. Kol Tregaskes pointed out this is the fourth Flash model in a row "with no Pro model." That is not a knock on 3.8 Flash itself, but it is the context: reporting suggests 3.5 Pro is being skipped entirely, so Flash is carrying flagship load while a Pro-tier model is missing from the lineup. If you are comparing across vendors, ChatGPT vs Gemini and my Gemini alternatives roundup are the two most useful starting points.

The verdict: should you switch from 3.7 Flash?

Because the price is identical, this is the only comparison that matters, and it is cleaner than most upgrade decisions.

If your workload is...The pickWhy
Long-horizon agentic coding, doc-heavy tasks, an overnight batch agentGemini 3.8 FlashGenuinely better on the benchmarks that matter here, at the same price; the extra thinking pays off
Latency-sensitive, human-in-the-loop (chat, live assist)Gemini 3.7 FlashThe 13.30s first-token cost hits a person directly; extra verbosity is pure overhead
Cost-first, high-volume, simple calls (classification, OCR, extraction)Gemini 3.7 FlashGoogle says so outright: it "remains fully supported for efficiency-first workloads"
An auditor to double-check a bigger model's workEither FlashCheap and fast enough that running it as a reviewer is a smart pattern

The verdict is that 3.8 Flash is a real upgrade for exactly one shape of work, and a downgrade in disguise for another. Google handing you both at the same price is not generosity, it is an admission that the newer model trades efficiency for diligence, and only you know which side of that trade your workload is on. One more thing to price in: the $0.75 rate is a promo that doubles to $1.50 on January 1, 2027. Anyone budgeting on today's number is budgeting on a discount.

What this means if you're putting AI on a support queue

Here is where I have to be straight, because I test these models specifically to answer the question support teams keep asking me: "which model should we point at our tickets?"

The honest answer is that the model is the wrong thing to be shopping for. A support workload is the near-perfect inverse of what 3.8 Flash is good at. Tickets are short-context, not long-horizon. Volume is high, so latency compounds, and a 13-second first-token delay times thousands of conversations is a real cost. And a raw model gives you none of the parts that actually make support automation work: it does not know your help centre, cannot read your past tickets, will not sit inside Zendesk or Freshdesk, and has no way to show you what it would have said before it says it to a customer.

I have watched confident-sounding bots quietly give wrong answers on live queues, which is the failure a benchmark score can never warn you about. That is why the thing I care about is not the model's Intelligence Index but whether the system around it can be simulated against your real tickets first. A model is infrastructure. What a support team actually needs is the employee, wired into the helpdesk with the context and the guardrails already in place. Model choice underneath that is a detail the platform should handle, not a decision you should be making by reading launch benchmarks.

Try eesel

If what you actually want is AI answering tickets rather than a model to wire up yourself, eesel is where I would start. eesel sells ready-to-work AI teammates rather than raw infrastructure, and the AI helpdesk teammate joins your existing queue in Zendesk, Freshdesk or Gorgias, trained on your help centre and your past tickets, in a few minutes.

The eesel AI helpdesk dashboard
The eesel AI helpdesk dashboard

The differentiator is the part model benchmarks cannot give you. Every eesel rollout gets simulated against your real historical tickets first, so you see what the AI would have said to customers you already served, before it says anything to a live one. It handles the unglamorous plumbing too, like ticket classification and support tagging, and it is free to try. The AI blog writer teammate runs on the same platform if content, not support, is your bottleneck.

On the numbers side, AI support cost savings has the maths, and agent versus human cost is the comparison finance teams actually ask for. And if you are still model-shopping rather than product-shopping, top AI agents and best AI agents for small business are the two roundups I would read next.

Frequently Asked Questions

Is Gemini 3.8 Flash worth switching to from 3.7 Flash?
For long-running agentic coding and document-heavy work, yes, it is a real step up at the same price. For latency-sensitive, high-volume jobs like a chat widget or ticket deflection, no. Google itself says to stay on 3.7 Flash for efficiency-first workloads, and the 13.30s time-to-first-token in this Gemini 3.8 Flash review is why.
How much does Gemini 3.8 Flash cost?
It is $0.75 per 1M input tokens and $3.75 per 1M output tokens (thinking tokens included) through Dec 31 2026, then it doubles to $1.50 / $7.50 on Jan 1 2027. That is identical to 3.7 Flash. The full breakdown, plus a calculator, is in my Gemini 3.8 Flash pricing post.
Why is Gemini 3.8 Flash so verbose?
By design. Google says it "works harder" by running extra reasoning steps and calling tools iteratively. Artificial Analysis measured 120M output tokens against a 71M median to run the same benchmark suite. Since output tokens include thinking tokens billed at the output rate, a model that thinks more also costs more per task, which is worth simulating before you commit, the way AI customer service automation should be tested against real volume.
Is Gemini 3.8 Flash good for a customer support chatbot?
It is the wrong shape for it. Support is short-context, high-volume, and latency-sensitive, and a model tuned for long-horizon coding that burns extra thinking tokens fights all three. For customer support automation you want a purpose-built AI helpdesk agent that is tested on your own tickets, not a raw coding model.
How does Gemini 3.8 Flash compare to Claude Opus 5?
On Artificial Analysis it scores 59 on the Intelligence Index, matching Opus 5 on medium effort, and it is cheaper and faster. But Opus 5 max scores 63 and uses roughly 4x fewer tokens for the same result, and it beats 3.8 Flash badly on Terminal-Bench 4.0 (51.8% vs 19.1%). See Gemini alternatives and Claude reviews for the wider field.

Share this article

Kurnia Kharisma Agung Samiadjie

Article by

Kurnia Kharisma Agung Samiadjie

Kurnia is a software engineer and writer at eesel AI with two years of SEO experience, writing about AI tools, helpdesk software, and customer support. He pairs a developer's understanding of how these products are built with search-driven research into what actually ranks and resonates with the people searching for them.

Related Posts

All posts →
Illustration of a fast-moving robot coding on a laptop while a person watches, representing Gemini 3.8 Flash
Trending

Gemini 3.8 Flash: what it is, honest benchmarks, and my review

Google shipped Gemini 3.8 Flash on September 2, 2026, three weeks after 3.7. Same price, better scores, and one line of fine print that changes the answer.

Alicia Kirana UtomoAlicia Kirana UtomoSep 3, 2026
Illustration of two people reviewing charts and speed dials around a Gemini spark, representing Gemini 3.8 Flash pricing
Trending

Gemini 3.8 Flash pricing: every rate, the hidden cost, and the catch

Gemini 3.8 Flash costs $0.75/$3.75 per 1M tokens, exactly what 3.7 Flash costs. But the sticker price hides a verbosity tax, and both numbers double on 1 January 2027.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieSep 8, 2026
Illustration of a developer and a colleague working with a fast AI coding agent
Trending

Gemini 3.7 Flash review: a great model that stopped being cheap

I put Google's Gemini 3.7 Flash against its own benchmarks and its own price list. It is fast and sharp, but it is no longer the cheap high-volume workhorse.

Rama Adi NugrahaRama Adi NugrahaAug 14, 2026
A person at a laptop beside a shield-shaped panel showing a tick, a question mark and a cross, with the Mistral mark on an orange background
Trending

Shieldstral review: a fast yes/no, and no reason why

Mistral's 3B open-weights safety classifier ties the 20B leader on its own text-safety chart and runs on one 16GB GPU. What it does not give you is a reason, or a hosted endpoint.

Alicia Kirana UtomoAlicia Kirana UtomoAug 18, 2026
A runner carrying a lightning bolt sprinting past a piggy bank, illustrating GLM-5.3 Flash speed and low cost
Trending

GLM-5.3 Flash review: frontier scores at flash cost

A hands-on GLM-5.3 Flash review: the benchmarks it actually posts, what its 4.5-cent-a-task price hides, where it breaks, and who should run it.

Kurnia Kharisma Agung SamiadjieKurnia Kharisma Agung SamiadjieAug 29, 2026
A reviewer looking at a verdict scorecard with two effort dials labelled low and max, beside the DeepSeek whale
Trending

DeepSeek V4 Flash review: one model, two personalities

A DeepSeek V4 Flash review built on the numbers both scoreboards publish. The cheap run and the smart run are the same weights, and that changes the verdict.

Riellvriany IndriawanRiellvriany IndriawanAug 4, 2026
DeepSeek V4 Flash pricing: what you'll actually be billed
Trending

DeepSeek V4 Flash pricing: what you'll actually be billed

DeepSeek V4 Flash lists at $0.14 in and $0.28 out per million tokens. Real users have posted blended rates under a cent. Here is what decides which one you get.

Alicia Kirana UtomoAlicia Kirana UtomoAug 4, 2026
DeepSeek V4 Flash: specs, pricing, and what it's really for
Trending

DeepSeek V4 Flash: specs, pricing, and what it's really for

DeepSeek V4 Flash costs $0.14 in and $0.28 out per million tokens, and it outscores DeepSeek's own expensive tier. Here's what the price card doesn't tell you.

Rama Adi NugrahaRama Adi NugrahaAug 4, 2026
Two people arm wrestling across a table while a third watches, illustrating a head-to-head model comparison
Trending

DeepSeek V4 Flash vs GPT-5.6: which one do you build on?

DeepSeek V4 Flash vs GPT-5.6 on August 2026 numbers. The real fight is Flash against Luna, intelligence is a tie, and the deciding factors are speed, vision and data.

Rama Adi NugrahaRama Adi NugrahaAug 4, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free