Gemini 3.8 Live Avatar review: is the talking AI face worth it?

Rama Adi Nugraha
Written by

Rama Adi Nugraha

Katelin Teen
Reviewed by

Katelin Teen

Last edited September 28, 2026

Expert Verified
A person talking to a lifelike lip-synced AI video avatar on a screen, illustrating a Gemini 3.8 Live Avatar review

How I reviewed this

A quick note on method, because it sets the basis for everything below. Gemini 3.8 Live Avatar is an enterprise feature with custom avatar creation behind an allowlist, so this isn't a "I ran it in production for six months" review. What I did was read Google's own model card, the developer guide, and the pricing table line by line, cross-check the launch claims against the named customer deployments, and pressure-test the whole thing against a question I care about more than looks: does a face make the answer any better?

I've spent the last few years building AI agents that sit on live support queues, and the thing that keeps me up at night was never how the bot looked. It's whether it's right. So when Google shipped a photoreal talking head for enterprise agents, my first question wasn't "does it look good" (it does), it was "does it move the number that matters" (it doesn't). Here's the full breakdown.

What you're actually getting

Google first previewed this at Google Cloud Next 2026 and shipped it to general availability the week after the base Gemini 3.8 Live and 3.8 Live Extended Thinking models launched. Worth getting the naming straight up front, because it's easy to trip on: Gemini 3.8 Live is the real-time conversation model, and Live Avatar is the animated talking-head video layer that sits on top of it. It's a feature, not a separate model, and per the DeepMind model card it's built on Gemini 3 Pro.

A verdict scorecard of where Gemini 3.8 Live Avatar shines versus where it stops short
A verdict scorecard of where Gemini 3.8 Live Avatar shines versus where it stops short

Strip away the marketing and the feature does five concrete things:

  • A real-time face. Precise lip-syncing, natural facial expressions, and fluid turn-taking, rendered as 24 FPS synchronized video output. The baseline 2.5 Flash Live model had no video output at all, so this is a genuine step, not a tweak.
  • It sees while it talks. The model processes audio and visual input at the same time, reading a camera feed or a screen share at one frame per second, so it can react to what you're pointing at mid-conversation.
  • 97 languages with mid-chat switching. Native multilingual speech-to-speech that adapts lip-sync and expressions across languages, switching mid-conversation with no visual drift. This is the standout capability, and I'll come back to it.
  • Background tool calls. Asynchronous function calling that fires tools and fetches data in the background while the dialogue keeps going, so the avatar isn't frozen mid-sentence waiting on an API.
  • Preset or custom avatars. A library of curated preset avatars anyone on Gemini Enterprise can use, plus custom avatars generated from a single reference image and audio sample, preserving likeness and brand styling.

Two settings are on by default and worth knowing about: affective dialogue (it reads prosody, emotion, and pauses, then adjusts its tone) and proactive audio (it filters out ambient noise and off-topic chatter). Every audio and video output also carries an imperceptible SynthID watermark, and custom avatar creation is gated behind allowlisting plus verification, which is Google being sensibly cautious about deepfakes.

The multilingual piece is the real headline

If I had to point at one thing that makes this worth watching, it's the language range. Ninety-seven languages, detected automatically, with the lip movements and expressions re-syncing when the conversation switches language mid-sentence. For a global support org that today staffs separate voice teams per region, that's a structurally different offer.

And Gemini's underlying language quality is genuinely strong, which is the part that makes the avatar layer interesting rather than gimmicky. One developer on the Hacker News thread for the 3.8 Live launch put it better than any benchmark could:

Hacker News

"My first language is Afrikaans, which is a somewhat niche language and hard to find teachers/conversation buddies outside South Africa... I've been using Gemini to live chat in Afrikaans... It is phenomenal at speaking the language - like, it really shocks my family members when they hear it."

That's the capability doing real work: not a demo language, a niche one, spoken well enough to surprise native speakers. Pair that quality with a face that lip-syncs correctly across the switch, and you can see why Equal AI is running it across 9 Indian languages at a reported 1M+ live calls per day.

What it actually costs

Here's where a review earns its keep, because "per token" pricing hides the real number. The avatar video is billed through the Gemini Enterprise pricing at $1.00 per 1M tokens, and the billing footnote is the important bit: avatar video output converts at 6,192 tokens per second of speech, and it's charged only while the avatar is actively speaking, not while it sits idle listening.

Do the math and that's about $0.0062 per second, or roughly $0.37 per minute of avatar talk time. On its own, the face is one of the cheaper lines on the bill.

A cost breakdown showing one minute of avatar talk time against the other Gemini 3.8 Live meters
A cost breakdown showing one minute of avatar talk time against the other Gemini 3.8 Live meters

The catch is that the avatar never runs alone. It sits on top of the full Live API meter stack, and here's every rate for the non-global endpoints:

MeterRate (per 1M tokens)Notes
Input text$0.75
Input video / image$1.00258 tokens per image
Input audio$3.0025 tokens per second
Output text (response + reasoning)$4.50reasoning tokens billed here
Output audio$12.00the expensive one
Output video (avatar)$1.006,192 tokens/sec, billed only while speaking

The line that dominates a spoken conversation is output audio at $12.00 per 1M, not the avatar video. So the honest framing is: turning on the avatar adds a modest, speech-time-only surcharge to a conversation that's already being billed for audio, reasoning, and inputs. There's also a structural tradeoff baked into the model card worth flagging: with Live Avatar enabled, the output token cap drops to 24K, down from 64K without it. And like any Live API session, past-turn tokens get re-processed and re-billed each new turn up to the context limit, so long conversations compound. If you're modelling this against a queue of thousands of daily interactions, build the estimate on audio and reasoning first, then add the avatar on top, not the other way round.

The catch: a face doesn't fix the answer

This is the whole review in one section. Everything above is real and impressive. None of it touches the thing that actually determines whether an AI support agent is worth deploying: is the answer correct?

A two-bar chart showing that the same avatar and voice score very differently depending on where the answer lives
A two-bar chart showing that the same avatar and voice score very differently depending on where the answer lives

The pattern I keep seeing across agent evaluations is a cliff. On tasks that resolve to a clean, structured tool call, top models are excellent. On tasks where the answer lives in a knowledge base article, an edge-case policy, or a half-documented exception, accuracy collapses. It's the same model, the same voice, the same face on both sides of that gap, and the score on the hard side is the one that decides whether customers get a right answer or a confident wrong one.

A lifelike avatar makes the wrong answer land harder, not softer. A blunt chatbot that says "I'm not sure, let me get a human" is annoying but honest. A warm, smiling, perfectly lip-synced face delivering an incorrect refund policy in flawless Portuguese is a worse outcome, because it's more persuasive. The presentation layer and the correctness layer are separate problems, and Google shipped an outstanding presentation layer.

That's not a knock on Gemini. It's a knock on treating a face as the finish line. The uncanny-valley risk is real too: a face that's 95% convincing can read as more unsettling than an obviously-synthetic one, and for some support moments (a frustrated customer, a billing dispute) a photoreal avatar is the wrong register entirely. Whether it helps or hurts depends on the interaction, and that's a design call, not a default.

Who this is actually for

After all that, here's who I'd tell to jump in and who I'd tell to wait.

Reach for it if you're building a voice-first or kiosk experience where the face is the product: an in-store assistant, a guided onboarding walkthrough, a conversational shopping helper. Cox Automotive and Autotrader are using it for exactly that, a shopping assistant with live screen-highlighting, and Marianne Johnson, their EVP and Chief Product Officer, is on record backing the approach. If you also need broad multilingual coverage and you can get on the enterprise allowlist for custom avatars, the fit is strong.

Wait, or look elsewhere, if your actual goal is resolving more support tickets correctly across email, chat, and your helpdesk. The avatar doesn't move that number, and you'd be paying a video meter and a token-cap tradeoff for a capability your core problem doesn't need. Text and voice channels resolve the vast majority of support volume, and none of them require a rendered face. This is also worth remembering if you're weighing it against a conversational AI platform built for support rather than a model you assemble yourself.

The deeper point: Gemini 3.8 Live Avatar is infrastructure. It's a building block you construct an agent from using Google's ADK and the Live API, with all the integration, grounding, testing, and guardrail work still on you. That's the right tool if you're a team building a custom product. It's the wrong tool if you want a support agent that's already a finished worker.

Try eesel for the answer, not the face

Here's the honest positioning, because it matters for a review: eesel doesn't build avatars, and it isn't a voice product. It's an AI helpdesk teammate that works your text support channels, and it's built around the exact problem the avatar leaves untouched: getting the answer right.

Where Gemini 3.8 Live Avatar is infrastructure you assemble, eesel is the teammate you hire. It plugs into your existing helpdesk (Zendesk, Freshdesk, and others), learns from your real past tickets and knowledge sources, and, crucially, lets you simulate it on thousands of your historical tickets before it ever replies to a live customer, so you see its accuracy and resolution rate on your own data first. It also reports exactly what it did, which tickets it handled, and where a human stepped in, so the correctness layer is measurable, not a leap of faith.

The eesel AI reports dashboard showing task volume, trigger events, and human approval usage
The eesel AI reports dashboard showing task volume, trigger events, and human approval usage

If a talking face is what your product needs, Gemini's is the best I've reviewed. If a correct answer is what your support queue needs, that's a different tool, and you can try eesel on your own tickets free.

Frequently Asked Questions

What is Gemini 3.8 Live Avatar?
It's a feature of Google's Gemini 3.8 Live model, generally available in Gemini Enterprise since September 24, 2026, that adds near real-time video generation to the speech-to-speech model. The result is a conversational agent with a lip-synced, expressive talking face that listens, watches a camera or screen, and speaks 97 languages. It's aimed at enterprise use like AI customer service and interactive product walkthroughs.
How much does Gemini 3.8 Live Avatar cost?
The avatar video output runs at 6,192 tokens per second of speech at $1.00 per 1M tokens, which works out to roughly $0.37 for every minute the avatar is actually talking. That sits on top of separate meters for output audio ($12.00/1M), output text and reasoning ($4.50/1M), and inputs, so the avatar layer is one of the cheaper lines on the bill. For the wider model family, see our Gemini pricing breakdown.
Can anyone create a custom Gemini Live Avatar?
No. You build a custom avatar from a single reference image and audio sample, but custom avatar creation is gated behind enterprise allowlisting and a verification step. Anyone on Gemini Enterprise can use the library of curated preset avatars without that allowlist, which is the faster path to trying it.
Is Gemini 3.8 Live Avatar good for customer service?
It's a strong presentation layer for voice-first and kiosk experiences, and the 97-language support is a real win for global teams. But a lifelike face doesn't change whether the answer is correct, which is the actual hard part of customer service automation. The value comes down to how well the underlying agent is grounded in your knowledge and systems.
How is Gemini Live Avatar different from an AI helpdesk agent?
Gemini Live Avatar is model infrastructure you build an agent with, using tools like Google's ADK. An AI helpdesk agent like eesel is a ready-to-work teammate that already plugs into your helpdesk, learns from your past tickets, and can be simulated before it goes live. One is a building block, the other is the finished worker.
What languages does Gemini 3.8 Live Avatar support?
97 languages, with automatic detection and lip-sync and expressions that adapt when the conversation switches language mid-sentence, with no visual drift. That multilingual range is genuinely impressive, though it's a separate thing from whether the answer is right in any of those languages.

Share this article

Rama Adi Nugraha

Article by

Rama Adi Nugraha

Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.

Related Posts

All posts →
A person talking to a lifelike AI video avatar on a screen, illustrating Gemini 3.8 Live Avatar
Trending

Gemini 3.8 Live Avatar: what it actually does for customer support

Google's Gemini 3.8 Live Avatar puts a lip-synced AI face on enterprise agents. I break down how it works, the real per-minute cost, and where it fits support.

Alicia Kirana UtomoAlicia Kirana UtomoSep 27, 2026
Illustration of a person talking to a live AI voice agent connected to several avatar faces
Trending

7 best Gemini 3.8 Live Avatar alternatives in 2026

The best Gemini 3.8 Live Avatar alternatives in 2026, from video-avatar platforms to voice-only APIs, with real pricing, latency, and an honest verdict on each.

Rama Adi NugrahaRama Adi NugrahaSep 28, 2026
Hand-drawn illustration of two people in a natural back-and-forth voice conversation, in OpenAI teal
Trending

GPT-Live-1 review: is OpenAI's full-duplex voice model worth it?

A hands-on GPT-Live-1 review: OpenAI's full-duplex voice model is the most natural I have used, but it delegates reasoning to a backend model and bills per minute. Here is the honest verdict.

Alicia Kirana UtomoAlicia Kirana UtomoSep 11, 2026
Illustration of a per-minute voice cost meter with sound waves, in OpenAI teal
Trending

GPT-Live-1 pricing: what the $0.05/min API really costs

GPT-Live-1 pricing is $0.05 per minute in the API, plus the backend model and harness you pair it with. Here is the full breakdown, the ChatGPT plans, and the real per-conversation cost.

Rama Adi NugrahaRama Adi NugrahaSep 11, 2026
Two people having a natural conversation with an AI voice assistant, sound waves flowing between them
Trending

GPT-Live-1: OpenAI's full-duplex voice model, explained

What GPT-Live-1 actually is: OpenAI's full-duplex voice model that listens and speaks at once, now in ChatGPT and the API at $0.05 per minute.

Alicia Kirana UtomoAlicia Kirana UtomoSep 11, 2026
A person holding a shield beside floating content cards and a checkmark panel, with the Mistral mark on an orange background
Trending

Shieldstral: accuracy is settled, packaging decides

Shieldstral ties a 20B model on text safety at 3B. The top four guard models sit inside 1.6 F1 points, so what actually picks your guard is hosting, licence, reasons, and how many calls one message costs.

Rama Adi NugrahaRama Adi NugrahaAug 18, 2026
Illustration of an AI assistant working across banking, fintech and insurance support
Trending

ChatGPT for financial services: what it does, security, and pricing (2026)

OpenAI just launched ChatGPT for financial services. Here is what it does, the security and compliance behind it, real pricing, and where it stops short.

Alicia Kirana UtomoAlicia Kirana UtomoSep 11, 2026
Illustration of Meta Muse, a personal AI agent, running errands inside a secure cloud computer
Trending

What is Meta Muse? Meta's personal AI agent, explained

Meta Muse is a personal AI agent that shops, books, and emails for you inside its own Secure VM. Here is what it does, how it works, and where it fits.

Alicia Kirana UtomoAlicia Kirana UtomoSep 9, 2026
Wonderful AI pricing breakdown illustration in deep electric blue
Trending

Wonderful AI pricing in 2026: what an enterprise AI OS really costs

Wonderful AI pricing is quote-only, with one public number: a $2.5M/year AWS listing. Here is what that buys, what it hides, and when to skip it.

Rama Adi NugrahaRama Adi NugrahaSep 9, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free