EmbeddingGemma 2: Google's open multimodal embedding model, explained

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 9, 2026

Expert Verified
Hand-drawn illustration of documents, images, audio and video flowing into a phone and coming out as one searchable list

What is EmbeddingGemma 2?

EmbeddingGemma 2 is an open embedding model from Google DeepMind, announced on October 6, 2026.

An embedding model does one job: it turns an input into a list of numbers (a vector) so that similar things end up close together. That is the engine under semantic search, clustering, and the retrieval half of RAG.

The first EmbeddingGemma, a 308M text-only model, came out in 2025 and passed 20 million downloads, per Google. Version 2 keeps the small, run-anywhere idea and turns it multimodal with four more input types. Per the model card, it maps text (including code), images, video and audio, or any mix of them, into a single 768-dimension space.

EmbeddingGemma 2 model card on Hugging Face, showing the Apache 2.0 license tag, 0.7B params in BF16, and 21,148 downloads last month, as taken from Hugging Face
EmbeddingGemma 2 model card on Hugging Face, showing the Apache 2.0 license tag, 0.7B params in BF16, and 21,148 downloads last month, as taken from Hugging Face

A quick note on where I'm coming from. I'm a software engineer at eesel and I build the AI agents that search a customer's help center before they answer a ticket. So when I read an embedding model launch, I'm less interested in the leaderboard and more in three questions: what does it find that the last one missed, what does it cost to run, and what breaks. I haven't put this model into production. My read comes from Google's model card and docs, the developer guide, and the first three days of hands-on reports on Hacker News.

EmbeddingGemma 2 at a glance

SpecEmbeddingGemma 2Source
ReleasedOctober 6, 2026Google blog
Parameters740M total: 270M text, 170M vision, 300M audioModel card
InputsText, code, images, video, audio, and interleaved mixesModel card
Output768 dims, truncatable to 512, 256, 128Developer guide
Context window8,192 tokens (v1: 2,048)Model card
Languages100+Gemma docs
RAM, quantized, Pixel 11 Pro~191MB text-only, ~567MB fullGoogle AI Edge
LicenseApache 2.0Gemma docs
PriceFree weights; you pay for computen/a
BaseGemma 4 decoder architectureGemma docs

How does EmbeddingGemma 2 work?

Under the hood, it is one small text model with two optional add-ons bolted to the front. The core is a 270M parameter model (a 130M transformer plus a 140M embedder) adapted from the Gemma 4 decoder. A 170M vision encoder handles images, video frames and visual documents like PDFs and slides. A 300M audio encoder takes raw speech and sound. Everything flows through the same backbone and comes out as one 768-number vector.

Diagram of EmbeddingGemma 2's modular design: text and code go straight into the 270M core, images and video pass through a 170M vision encoder, audio through a 300M audio encoder, and similar items land close together in one space, as taken from Google's developer guide
Diagram of EmbeddingGemma 2's modular design: text and code go straight into the 270M core, images and video pass through a 170M vision encoder, audio through a 300M audio encoder, and similar items land close together in one space, as taken from Google's developer guide

The clever part is that the add-ons are optional at load time. You set vision_config or audio_config to None and that encoder is never loaded, which saves both disk and memory. Per Google's developer guide, all four setups load from the same checkpoint and share one vector space. A query embedded with the 270M text-only setup can match documents embedded with the full model.

Hand-drawn chart of the four ways to load EmbeddingGemma 2: text only at 270M and about 191MB RAM, text plus vision at 440M, text plus audio at 570M, and the full 740M model at about 567MB RAM, all feeding one shared 768-dimension space
Hand-drawn chart of the four ways to load EmbeddingGemma 2: text only at 270M and about 191MB RAM, text plus vision at 440M, text plus audio at 570M, and the full 740M model at about 567MB RAM, all feeding one shared 768-dimension space

That has a practical upside I like a lot. You can start text-only and add images later without recomputing a single vector you already have, the guide says. If you have ever re-embedded a million documents because you changed models, you know how much that is worth.

Inputs share the 8,192-token context window at fixed rates, counted in tokens. One image costs 280 tokens by default, a video frame costs 140, and a second of audio costs 25. So a single call can hold about 29 images, 58 video frames, or 327 seconds of audio, per the model card.

Table of EmbeddingGemma 2 token costs per modality: 1 token per subword up to 8,192 tokens, 280 tokens per image for about 29 images, 140 per video frame for about 58 frames, and 25 per second of audio for about 327 seconds, as taken from Google's developer guide
Table of EmbeddingGemma 2 token costs per modality: 1 token per subword up to 8,192 tokens, 280 tokens per image for about 29 images, 140 per video frame for about 58 frames, and 25 per second of audio for about 327 seconds, as taken from Google's developer guide

A few mechanics worth knowing before you write code:

  • Video is sampled at 1 frame per second by default, and audio has to be 16 kHz mono.
  • Image detail is a dial. The multimodal docs list five budgets per image: 70, 140, 280, 560 or 1,120 tokens. More tokens mean finer detail and slower runs.
  • You can interleave. One input can mix text with <|image|>, <|video|> and <|audio|> placeholders, like a product page with a photo and a demo clip, and it still comes out as a single vector.
  • Text gets a task prefix, media does not. Queries look like task: search result | query: ... and documents like title: ... | text: .... Skipping the prefix still works but costs precision.

What changed from EmbeddingGemma 1?

Here's where I'd slow down, because the launch framing and the numbers tell slightly different stories. Google's announcement says v2 matches v1's multilingual text performance and jumps on code. That's true. But the model card also carries MTEB English, and on that one v2 is a little behind.

Hand-drawn scorecard comparing EmbeddingGemma 1 and 2: English text 69.67 to 68.46, multilingual text 61.15 to 61.36, code search 68.76 to 78.68, and context window 2K to 8K
Hand-drawn scorecard comparing EmbeddingGemma 1 and 2: English text 69.67 to 68.46, multilingual text 61.15 to 61.36, code search 68.76 to 78.68, and context window 2K to 8K
Benchmark (768 dims)EmbeddingGemma 1EmbeddingGemma 2Change
MTEB English v269.6768.46-1.21
MTEB Multilingual v261.1561.36+0.21
MTEB Code v168.7678.68+9.92
Context window2,048 tokens8,192 tokens4x
InputsTextText, code, image, video, audio+4
LicenseGemma terms, gated downloadApache 2.0Open

The v1 numbers come from the EmbeddingGemma 300M card and the v2 numbers from the v2 card. Early hands-on testers landed in the same place:

Hacker News

"for text, benchmarks are identical to the first EmbeddingGemma. but you can use this new one and enable/disable what you don't need. can keep only text for ex."

So my read is simple. If you only search English text, v2 is a sideways move, and the upgrade is worth it only if the 4x longer context helps you (fewer, bigger chunks) or the Apache license unblocks a product. If you search code, or anything that is not text, it is a big step. For a help center that is mostly articles, my take on RAG vs fine-tuning matters more than this upgrade.

What the new benchmarks say

For the new input types there is no v1 to compare against, so the model card scores stand alone. Google calls it "best-in-class" among sub-1B multimodal embedders on MTEB Code and MAEB, the audio benchmark. That is a vendor claim with no head-to-head table on the card, so treat it as a starting point.

InputBenchmarkScore
ImageMIEB lite64.64
ImageMMEB v2 Image (Hit@1)57.28
Visual documentsMMEB v2 VisDoc (NDCG@5)67.84
VideoMMEB v2 Video (Hit@1)50.67
AudioMSEB Retrieval (MRR@10)69.54
AudioMAEB49.39

Two things I'd flag from the HN thread. One commenter asked why it isn't compared to SigLIP 2, Google's own image-text model. Another got it running locally and found the zero-shot classification example from Google's own demo misfired:

Hacker News

"I got this running locally, and funnily enough the exact example they have for classifying "Cancel my flight and refund my credit card immediately." failed. It said the request does not involve a payment, charge, or refund, with a p(true) of 0.22 (where true means it is financial)."

That one stings a little for anyone eyeing it for ticket routing, since "refund or not" is exactly the kind of label a support team would want. One data point is not a verdict, but it is a good reason to test on your own tickets before wiring it to anything.

There's also a quieter gotcha in the multimodal docs: similarity scores are not calibrated across modalities. An unrelated painting scored 0.585 against a Golden Gate photo, higher than unrelated text prompts did. Compare rank within a search, not raw scores between images and text.

How fast and small is it, really?

This is where EmbeddingGemma 2 earns the "on-device" label. Google's AI Edge team published per-device numbers for embedding one image at a 70-token budget:

DeviceChipLatency per imageImages/sec
iPhone 18 ProCPU191 ms5.2
iPhone 18 ProGPU69.8 ms14.3
Pixel 11 Pro XLNPU48.9 ms20.6
MacBook M5 ProGPU37.3 ms26.9
Dell XPS 16NPU49.8 ms20.1
Raspberry Pi 5CPU1,761 ms0.6

Real-world numbers from a laptop look similar. One HN commenter shared throughput from an M3 Pro:

Hacker News

"Text: 78 embeddings/second on small texts - Image: 4 embeddings/second - Audio: 6 embeddings/second for 30 second chunks (which is how the encoder works) - Video: 0.2 embeddings/second per minute of video"

Note the gap. Video at 0.2 per minute of footage means an hour of video takes about five minutes to index on that laptop. Fine for a library of product demos. Too slow for live, high-volume video.

Storage is the other lever. The model uses Matryoshka training, so you can chop vectors from 768 numbers down to 512, 256 or 128 and keep most of the quality. Per the developer guide, a million 768-dim vectors take about 1.5GB in bfloat16, and about 250MB at 128 dims.

Hand-drawn storage boxes shrinking from 768 dims (1M vectors in 1.5GB) to 256 dims (3x smaller, about 95% quality on images, video and audio) to 128 dims (6x smaller, about 90% on text, about 75% on images, video and audio), with 256 marked as the safe default
Hand-drawn storage boxes shrinking from 768 dims (1M vectors in 1.5GB) to 256 dims (3x smaller, about 95% quality on images, video and audio) to 128 dims (6x smaller, about 90% on text, about 75% on images, video and audio), with 256 marked as the safe default

The catch is that media degrades faster than text when you squeeze. On the model card, MMEB v2 drops from 59.01 at 768 dims to 45.65 at 128. For a mixed index, I'd stop at 256 dims and only go to 128 for text-only shortlists you re-rank afterwards. Two rules from the card that will bite if you skip them: re-normalize vectors after you truncate, and never run it in float16. The activations overflow and you get NaNs or silently worse vectors with no error.

Which EmbeddingGemma 2 setup should you load?

Since you pay memory only for the encoders you load, the first decision is what your index actually contains. Pick the closest match:

What goes into your search index?

Tap an option to see the setup I'd start with.

Text only, 270M, about 191MB RAM. Load with vision_config=None, audio_config=None. Honest note: for English-only search, v1 scored slightly higher (69.67 vs 68.46). Switch if you need the 8K context or Apache 2.0, not for accuracy. 256 dims is a safe storage cut.
Text only, 270M. This is where v2 shines: 78.68 on MTEB Code vs 68.76 for v1. Use the CodeRetrieval prompt for queries and put the filename in the document title. Stay at 512 or 768 dims; code loses quality fastest when truncated.
Text plus vision, 440M. Load with audio_config=None. Vision covers images, PDFs, slides and charts (MMEB v2 VisDoc 67.84). Raise the image budget to 560 or 1,120 tokens if your screenshots have small text. Keep 256 dims or more.
Text plus audio, 570M. Load with vision_config=None. Resample audio to 16 kHz mono and chunk it; one call holds about 327 seconds. It matches speech content, not just "this is speech", but test on your own calls first.
Full model, 740M, about 567MB RAM. Load with no config changes. Video is sampled at 1 frame per second; expect roughly 5 minutes to index an hour of footage on a recent laptop. Stay at 768 or 512 dims, since video and audio lose the most when truncated.

How do you run EmbeddingGemma 2?

Weights are on Hugging Face and Kaggle (new to the platform? start with my Hugging Face explainer), and day-one support is wide: sentence-transformers (v6.1.0 or later), Transformers, vLLM, SGLang, MLX, llama.cpp GGUF, Ollama, LM Studio, and Google's own LiteRT for phones. The shortest path, straight from the model card:

Python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

query_emb = model.encode("What causes the northern lights?", prompt_name="SearchQuery")
doc_emb = model.encode("The northern lights are caused by charged particles from the sun.", prompt_name="Document")
print(model.similarity(query_emb, doc_emb))

If you'd rather not touch Python, the Ollama build plugs into most local RAG stacks, and frameworks like LangChain and LlamaIndex will treat it like any other embedder.

For a worked support example, see my support RAG pipeline walkthrough or this Zendesk semantic search guide. You still need somewhere to store the vectors; Qdrant shipped a guide on launch day, and my post on RAG vs vector databases covers how to choose.

On phones, Google is bringing it to Android as a managed ML Kit service "in the coming weeks", with NPU acceleration. Until then, the AI Edge Gallery app is the easiest way to feel it. Its Instant Media Search demo searches your camera roll as you type, offline, and the Video Moments Finder jumps to timestamps for queries like "dog catching a frisbee" with no transcript.

Should you self-host it or use Gemini Embedding 2?

This is the second reframe I'd offer. EmbeddingGemma 2 is free to download, and people often assume that makes it the cheap option. For most teams, it isn't the cheaper one by much, because Google's hosted sibling is already very cheap.

Gemini Embedding 2 is Google's paid multimodal embedding API, and it is the only embedding model on the current Gemini API price list. (If you're weighing Google's API more broadly, see OpenAI API vs Gemini API.) My Gemini pricing post covers the rest; here are the embedding rates:

InputStandard (per 1M tokens)BatchPer unit (standard)
Text$0.20$0.10n/a
Image$0.45$0.225$0.00012 per image
Audio$6.50$3.25$0.00016 per second
Video$12.00$6.00$0.00079 per frame

Now a worked example for a mid-size support team. Say 2,000 help articles at about 800 tokens each (1.6M tokens), 10,000 product screenshots, and 100 hours of recorded calls:

  • Articles: 1.6M tokens x $0.20 = $0.32
  • Screenshots: 10,000 x $0.00012 = $1.20
  • Calls: 360,000 seconds x $0.00016 = $57.60

That is under $60 to embed the whole library once at standard rates, and half that in batch. A rented GPU at the $1 to $2 an hour one HN commenter quoted would run longer than that just to get going. So I'd frame the choice this way:

  • Use Gemini Embedding 2 if your data can go to Google's API and you want zero setup. Just note that the free tier's data can be used to improve Google's products; the paid tier's is not.
  • Use EmbeddingGemma 2 if the data cannot leave the device or your own servers (health, legal, internal support), if you need it offline, or if you embed queries on a phone and want no round trip. Those are real reasons. Cost alone usually isn't.

Simon Willison put the case for open embedders well:

Hacker News

"I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license. For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model."

His point is about lock-in more than price. Every vector in your index is tied to the model that made it. If a hosted model is retired, you re-embed everything. With open weights, the model you indexed with stays available as long as you keep a copy.

Why support teams should care about the multimodal part

Here's the angle I haven't seen in the launch coverage. Most writing about RAG for customer service assumes the knowledge is text. In practice, a lot of it isn't, and it comes up on eesel's sales calls constantly.

A DTC supplements brand doing about 7,000 tickets a month said their answers lived in ClickUp SOPs, untranscribed Loom videos, and old macros. A public-sector IT firm handling around 3,000 complex tickets a month runs on Freshdesk with image-heavy technical docs. Technical-product teams keep asking for AI that can read the diagrams inside their PDFs and the photos customers attach to tickets. Text-only embedders can't see any of that unless someone transcribes or captions it first, which is why so many support knowledge base tools still treat media as an attachment, not an answer.

Hand-drawn diagram of support knowledge that is not plain text, a help article, product screenshot, Loom walkthrough, PDF wiring diagram and call recording, all flowing into one index, so a search for "LED blinks red after reset" returns a screenshot and a video moment at 1:42
Hand-drawn diagram of support knowledge that is not plain text, a help article, product screenshot, Loom walkthrough, PDF wiring diagram and call recording, all flowing into one index, so a search for "LED blinks red after reset" returns a screenshot and a video moment at 1:42

One shared space is the fix for that. Index the screenshot, the walkthrough and the article together, and a text query like "LED blinks red after reset" can pull back the right frame of the right video. One HN reader, a lawyer, described the same idea for case files:

Hacker News

"I could imagine using it to search a case file for "undamaged roof before Hurricane Katrina" and "damaged roof after Hurricane Katrina" and being able to locate both deposition testimony and pertinent photographs in the body of evidence."

But retrieval is only half of a support bot, and in my experience it is not the half that hurts. Earlier this year a few paying eesel customers, including a Danish solar-energy provider, saw their bot invent answers when the knowledge base search came back empty. It filled the gap from the model's training data, a textbook case of AI hallucination, and sent made-up subscription details to real customers.

A better embedder would not have fixed that. A hard "no match, hand it to a human" rule did, plus simulating the bot against past tickets before go-live. If you build on EmbeddingGemma 2, plan for that layer from day one; my guide on stopping AI hallucinations covers the patterns.

Who should use EmbeddingGemma 2?

You areMy take
Building search over a codebase or coding-agent memoryUse it. +9.92 points on MTEB Code is the clearest win in the release.
Searching photos, screenshots, PDFs, audio or videoUse it, at 256 dims or more, and test on your own files.
Shipping on-device or offline search in an appUse it. 191MB for text is small enough for a phone, and Android ML Kit support is coming.
Already on EmbeddingGemma 1 for English text onlyStay put unless you need 8K context or the Apache license.
Running support search in the cloud with no data limitsGemini Embedding 2 is less work for about the same money.
Wanting a working support bot, not a pipelineSkip the build. The model is one part of five.

For comparisons with other options, my posts on Cohere Embed 5 and the OpenAI embeddings API cover the main hosted alternatives, and small language models goes deeper on the on-device trend this model belongs to.

For the wider field, my roundup of knowledge retrieval tools compares the finished products.

The verdict

EmbeddingGemma 2 is a good, honest release. The modular design is smart, the Apache license removes a real blocker, and code and multimodal search take a clear step forward in a package small enough for a phone. If your data is not just text, it is the first open model I'd try.

Just read past the headline. For English text it is level with the model it replaces, the "best-in-class" claims are Google's own, and for most cloud teams the hosted API costs about the same as running it yourself. Pick it for privacy, offline use, code, or media, and you'll be happy. Pick it expecting a big jump on plain-text search, and you won't see one.

Try eesel

If you came here because you are thinking about building a support bot on a fresh embedding model, here's the shortcut: the embedder finds passages, but someone still has to read them, write the reply, and know when to stay quiet. eesel's AI helpdesk teammate does that job inside Zendesk, Freshdesk or Gorgias. It learns from your help center, your docs in places like Confluence and Google Docs, and your past tickets, then runs a simulation over hundreds of your real tickets so you see its answers before a customer does. When search comes back empty, it hands the ticket to a person instead of guessing.

eesel's Skills list showing Simulation, which runs the agent against multiple past tickets and returns a scored performance report, alongside Support Analytics and Self Review
eesel's Skills list showing Simulation, which runs the agent against multiple past tickets and returns a scored performance report, alongside Support Analytics and Self Review

There's a free plan with 100 credits and no card, and paid plans are on the pricing page. Try eesel and replay last month's tickets before you build anything.

Frequently Asked Questions

What is EmbeddingGemma 2?
EmbeddingGemma 2 is Google DeepMind's open embedding model, released on October 6, 2026. It turns text, code, images, audio and video into vectors in one shared 768-dimension space, so a text query can find a matching photo or video moment. It has 740M parameters in total and is built on Gemma 4.
Is EmbeddingGemma 2 free to use commercially?
Yes. EmbeddingGemma 2 ships under the Apache 2.0 license, which allows commercial use, fine-tuning and redistribution. That is a change from the first EmbeddingGemma, which used the Gemma terms and a gated download. Your only cost is the hardware you run it on, as with other Hugging Face open models.
How much RAM does EmbeddingGemma 2 need?
Google measured about 191MB of active RAM for the text-only setup and about 567MB for the full multimodal model, both quantized on a Pixel 11 Pro. You can load only the encoders you need: 270M parameters for text, 440M with vision, 570M with audio, or 740M for everything.
Can I use EmbeddingGemma 2 for a support chatbot?
You can use it as the search layer in a RAG pipeline over your help center, but it only finds passages; it does not write replies. You still need an LLM, a vector database, and a fallback for when nothing matches. A ready-made option like eesel's AI helpdesk teammate handles that whole loop inside your helpdesk.
How does EmbeddingGemma 2 compare to Gemini Embedding 2?
Gemini Embedding 2 is Google's hosted multimodal embedding API, priced at $0.20 per 1M text tokens on the paid tier. EmbeddingGemma 2 is the open, self-hosted sibling that runs offline on a phone or laptop. Pick the API for zero setup, and the open model when data cannot leave your device. See my Gemini pricing breakdown for the rest of the API.
Do I need to re-embed my data to switch to EmbeddingGemma 2?
Yes, if you are moving from another model, including EmbeddingGemma 1. Vectors from different models are not comparable, so the whole index has to be rebuilt. Within EmbeddingGemma 2 you do not: adding the vision or audio encoder later keeps your existing text vectors valid, per Google's developer guide.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Hand-drawn illustration of three people at a laptop looking up at an idea, in Reflection AI's dark green
Trending

Reflection AI Beam: what the 501B open-weight model is, and what you can use today

Reflection AI Beam is a 501B open-weight model with 23B active parameters. Here is what it scores, what is actually usable today, and who should wait.

KiraKiraOct 8, 2026
Cloudflare Clef hero banner in Cloudflare orange, two people looking at a decision model connected to users, websites, devices and a list of options
Trending

What is Cloudflare Clef? Cloudflare's decision model, explained

Cloudflare Clef explained: what the decision model does, how Clef differs from Clef-flash, how to run it on Workers AI or Ollama, and where it fits in ticket routing.

KiraKiraOct 7, 2026
Gemini Omni Flash review hero banner in Google blue
Trending

Gemini Omni Flash review: Google's fast, cheap AI video model

A hands-on Gemini Omni Flash review: what Google's new AI video model does, what it costs at $0.10/sec, where it lags Seedance, and who should use it.

KiraKiraJul 11, 2026
Hand-drawn illustration of a reviewer with a clipboard scorecard studying a glowing crystal model inside a roped-off glass display case, for a Gemini 4 Argon review
Trending

Gemini 4 Argon review: a great model you can't use yet

My Gemini 4 Argon review: the lowest hallucination rate of any top model and Astra-level intelligence, but a 404 on the API, heavy token use, and a price that doubles.

Rama AdiRama AdiOct 1, 2026
Gemini 3.5 Pro review hero banner in Google blue
Trending

Gemini 3.5 Pro review: the honest state of Google's flagship

An honest Gemini 3.5 Pro review: it isn't out yet. Here's what Google has confirmed, why it's late, the benchmarks that do exist, and what to use today.

Kurnia KharismaKurnia KharismaJul 21, 2026
GPT-5.6 versus Gemini 3 comparison hero illustration, two AI model families balanced against each other
Trending

GPT-5.6 vs Gemini 3: which AI model wins in 2026?

GPT-5.6 vs Gemini 3 compared: Sol, Terra and Luna against Gemini 3.5 Flash and 3.1 Pro on pricing, benchmarks, context, and which fits AI support agents.

Rama AdiRama AdiJul 10, 2026
A person holding a shield beside floating content cards and a checkmark panel, with the Mistral mark on an orange background
Trending

Shieldstral: accuracy is settled, packaging decides

Shieldstral ties a 20B model on text safety at 3B. The top four guard models sit inside 1.6 F1 points, so what actually picks your guard is hosting, licence, reasons, and how many calls one message costs.

Rama AdiRama AdiAug 18, 2026
Hand-drawn illustration of two people looking at a price tag and a speed gauge next to the Mistral logo
Trending

Mistral Large 4 pricing: API rates, service tiers, and the real cost per task

Mistral Large 4 pricing is $0.68 in and $2.09 out per 1M tokens on sale. Regional and Priority tiers are priced off the $1.36 / $4.18 list price, not the sale.

Rama AdiRama AdiOct 8, 2026
Hand-drawn illustration of three people around a table discussing a stack of server blocks topped with the Mistral logo
Trending

Mistral Large 4: specs, pricing, benchmarks, and who should use it

Mistral Large 4 is a 1T-parameter open-weight preview at $0.68/$2.09 per 1M tokens. A huge jump for Mistral, a mid-pack score next to rivals, and a few real niches.

KiraKiraOct 8, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free