Cohere Embed 5: Pro vs Fast, benchmarks, pricing, and how to use it

Kira
Written by

Kira

Katelin Teen
Reviewed by

Katelin Teen

Last edited October 1, 2026

Expert Verified
Two people searching a document index powered by Cohere Embed 5 embeddings

What Cohere Embed 5 actually is

Cohere is the enterprise AI company behind the Command models, Rerank, and the Parse 5 document parser. An embedding model is the quiet part of any search or RAG system, the piece that turns a chunk of text (or an image) into a list of numbers, and chunks that mean similar things then end up sitting close together. When a user asks something, that question also gets turned into numbers, then the system goes and fetches whichever chunks are nearest.

When the embedding model is weak, the right answer simply never gets fetched, and clever prompting is not going to fix that for you. It is the same layer that powers AI enterprise search.

Embed 5 takes over from Embed 4 as the flagship of Cohere, and it comes in two tiers: Pro, "optimized for maximum quality across multimodal, multilingual, financial, code, and parsed-document retrieval," and Fast, which "brings highly competitive performance to latency- and cost-sensitive workloads."

The Cohere Embed product page announcing Embed 5 in Pro and Fast tiers, as taken from Cohere

I build AI agents for a living at eesel, and retrieval is the layer where most of my debugging time goes, so these specs were the first ones I went looking for. Here they are side by side, pulled from Cohere's model docs and launch post:

Embed 5 ProEmbed 5 FastEmbed 4 (previous)
API model IDembed-v5.0-proembed-v5.0-fastembed-v4.0
Text price (per 1M tokens)$0.12$0.08$0.12
Image price (per 1M tokens)$0.40$0.40$0.47
Context length128K tokens128K tokens128K tokens
InputsText, images, fused text + imageText, images, fused text + imageText, images, mixed
Output dimensions256 to 2048 (default 2048)256 to 2048 (default 2048)256 to 1536 (default 1536)
Formatsfloat, int8, binaryfloat, int8, binaryfloat, int8, binary
Languages100+100+Multilingual
Best forOffline indexing, quality-critical searchLive queries, agent loops, high volume

The Embed 4 rates come from Microsoft Foundry's Cohere pricing, since Cohere's own pricing page no longer lists Embed 4. If you are coming from OpenAI's 3-series, note that its API pricing still tops out at $0.13 for text-embedding-3-large. In the Cohere table there are two small details that people tend to miss. The default output went up from 1536 to 2048 dimensions, which is something that matters for storage (more on that below), and also image input got cheaper, from $0.47 to $0.40 per 1M tokens.

How Embed 5 works: index with Pro, query with Fast

This is the feature I would build around, and it is the one with the least hype attached. Most embedding families, including the OpenAI embeddings API, make you pick one model and live with it, because vectors from two different models are not comparable. Cohere trained Pro and Fast into a single shared space, so a query embedded with Fast can be matched directly against documents embedded with Pro.

The reason this matters is that indexing and querying want opposite things. Indexing happens once (or on every document update) and nobody sits there waiting on it, so here you want the best quality money can buy. Querying is a different story, it happens on every single user request and inside your latency budget, which means speed is what counts. Having one shared space lets you choose each side on its own.

Diagram of the index-with-Pro, query-with-Fast pattern: documents go through Embed 5 Pro, live queries go through Embed 5 Fast, both land in one shared index at 98.4% of all-Pro quality
Diagram of the index-with-Pro, query-with-Fast pattern: documents go through Embed 5 Pro, live queries go through Embed 5 Fast, both land in one shared index at 98.4% of all-Pro quality

Cohere ran every pairing across 40 development datasets, then normalized the scores in a way that Pro documents plus Pro queries equals 100:

Mean retrieval qualityDocuments indexed with FastDocuments indexed with Pro
Queries with Fast96.698.4
Queries with Pro97.3100

So Pro index plus Fast queries lands at 98.4, a 1.6% loss, while Fast does the query-time work at roughly 2.4x the throughput of Pro (377.3 versus 159.7 documents per second on Cohere's test). Cohere itself recommends exactly this pattern, and in the footnote there is the catch worth knowing about: both sides have to use the same output dimension. The pairing still holds when you add Matryoshka truncation and int8 quantization, so a compressed index works with it too.

Bar chart of document throughput: Embed 5 Fast processes 377.3 documents per second versus 159.7 for Embed 5 Pro, as taken from Cohere
Bar chart of document throughput: Embed 5 Fast processes 377.3 documents per second versus 159.7 for Embed 5 Pro, as taken from Cohere

A few API details from the Embed reference that shape how you call it:

  • input_type is required. Use search_document when indexing and search_query at query time. There are also classification and clustering modes.
  • 96 inputs per call. Each input can mix text and image parts, with a 20MB total payload cap.
  • Truncation defaults to END. With a 128K context it is rare that you hit it, but if you prefer getting an error over silently losing the end of a long document, set truncate to NONE.
  • Rate limits are per input, not per request: 2,000 text inputs a minute on both trial and production keys, and 5 versus 400 image inputs a minute, per the rate limits page. Trial keys are free but capped at 1,000 calls a month and not allowed for production use.

The benchmarks: strong, but read the fine print

The numbers Cohere published are good ones. Below is the headline ViDoRe V3 table, it covers enterprise documents across eight domains:

ModelViDoRe V3 averageText price per 1M tokens
Cohere Embed 5 Pro85.8$0.12
Cohere Embed 5 Fast84.5$0.08
Voyage 4 Large83.7$0.12
Gemini Embedding 283.2$0.20
Cohere Embed 477.0$0.12
OpenAI text-embedding-3-large75.5$0.13
Jina Embeddings v5 Text Small74.5$0.05

On the same test Pro gains 8.8 points over Embed 4, and the biggest jumps are on HR (+11.4) and on industrial documents (+10.3). It is also the leader of Cohere's parsed-PDF suite at 84.8. For financial retrieval it ranks first on FinanceBench (80.1) and FinQA (90.0), plus ViDoRe V3 Finance (85.0), with Fast coming second on each of them.

Grouped bar charts of parsed-PDF retrieval quality across seven embedding models, with Embed 5 Pro highest on the average and on RepairBench, CoFiF, FinanceBench, and MPQMA, as taken from Cohere
Grouped bar charts of parsed-PDF retrieval quality across seven embedding models, with Embed 5 Pro highest on the average and on RepairBench, CoFiF, FinanceBench, and MPQMA, as taken from Cohere

For me the most impressive result is actually Fast. Put against other compact models on ViDoRe V3 it scores 84.5, where Voyage 4 Nano gets 77.6 and Qwen3-VL-Embedding gets 64.2. The legend on Cohere's own chart (below) lists Fast at about 1B parameters, 500M text plus 500M vision, and the launch post says it beats Qwen3-VL-Embedding-2B by about 20 points despite being roughly half its size.

Bar chart of ViDoRe V3 retrieval quality for compact models: Embed 5 Fast 84.5, Voyage 4 Nano 77.6, Jina Embeddings v5 Text Small 74.5, Perplexity pplx-embed-v1 74.4, Microsoft Harrier OSS v1 73.4, Qwen3-VL-Embedding 64.2, as taken from Cohere
Bar chart of ViDoRe V3 retrieval quality for compact models: Embed 5 Fast 84.5, Voyage 4 Nano 77.6, Jina Embeddings v5 Text Small 74.5, Perplexity pplx-embed-v1 74.4, Microsoft Harrier OSS v1 73.4, Qwen3-VL-Embedding 64.2, as taken from Cohere

Now the fine print, because a fair read needs it.

It is a new metric, run by the vendor. Embed 5 is the first model family scored with RCP-nDCG@10, a method Cohere published the same day. Cohere's own footnote says it evaluates models by reordering a fixed candidate set, so "scores therefore reflect reranking quality rather than first-stage retrieval performance." It is a reasonable way to measure, and the code is public, but it is not the same number you would get by running a plain top-10 search over your whole index. Also some of the datasets are internal ones, like the "High Finance" set which Cohere annotated by itself.

Gemini Embedding 2 wins most non-European languages. Pro leads the European set (77 average across German, French, Spanish, Italian, and Russian). But in Cohere's own ten-language table, Gemini Embedding 2 beats Pro on 9 of the 10: Japanese, Korean, Arabic, Farsi, Hindi, Bengali, Telugu, Indonesian, and Thai. Pro only edges it on Chinese (82 versus 81). Telugu is the widest gap, at 91 versus 80.

Scorecard showing where Embed 5 Pro leads (ViDoRe V3 85.8, parsed PDFs 84.8, European languages 77) and where Gemini Embedding 2 leads (Telugu 91 vs 80, Bengali 89 vs 83, 9 of 10 Asian and Indic languages)
Scorecard showing where Embed 5 Pro leads (ViDoRe V3 85.8, parsed PDFs 84.8, European languages 77) and where Gemini Embedding 2 leads (Telugu 91 vs 80, Bengali 89 vs 83, 9 of 10 Asian and Indic languages)

If your support queue or document base leans heavily on Hindi, Thai or Bengali, that table is the most useful thing in the whole launch, and credit to Cohere for printing it at all. For a team serving a multilingual knowledge base across Europe and English, Pro is the stronger pick on these numbers.

Community reaction is still early and thin, which is what you would expect one day after a launch. The most useful voice from a practitioner I found is an older one, and it is about Cohere's embeddings in general rather than Embed 5 itself:

Hacker News

"My experience with Cohere and interacting with their sales engineers has been boring, I say that is the most flattering way possible. Embeddings are a core service at this point like VMs and DBs. They just need to work and work well and thats what they're selling."

That matches the pitch here pretty well. Embed 5 is not trying to be exciting, what it tries to be is the boring and dependable layer that sits under your search.

Cohere Embed 5 pricing

Billing for Embed 5 is per input token, and there is no output charge. Here is everything Cohere publishes, from the pricing page and the launch post:

OptionEmbed 5 ProEmbed 5 FastBilling unit
Cohere API, text$0.12 per 1M tokens$0.08 per 1M tokensInput tokens
Cohere API, images$0.40 per 1M tokens$0.40 per 1M tokensImage tokens
Trial keyFree, 1,000 calls a monthFree, 1,000 calls a monthNot for production use
Model Vault Small$3.00/hour or $2,000/month$3.00/hour or $2,000/monthPer dedicated instance
Model Vault Medium$5.00/hour or $3,250/month$5.00/hour or $3,250/monthPer dedicated instance
Amazon SageMaker$2.39 to $8.48 per host-hour$2.39 to $3.36 per host-hourSoftware fee plus AWS instance cost
Microsoft FoundryNot yet published (preview)Not yet published (preview)

The SageMaker rates come from the AWS Marketplace listings for Embed 5 Pro, with a matching listing for Fast. Cohere bills at the end of each month, or sooner once you hit $250 outstanding. For the rest of the catalog (Rerank, Parse, Command), see my full Cohere pricing guide. If you are also parsing PDFs, the Parse 5 pricing breakdown covers that meter.

Three billing notes I would want to know about before setting a budget:

  1. Image token counts are not documented. The price is per 1M image tokens, but Cohere publishes no tokens-per-image formula. The API response reports images as a count ("images": 1), so run a small test batch and read the bill before you embed a million page images.
  2. Model Vault only pays off at very high volume. A Small instance at $2,000 a month equals about 16.7 billion Pro tokens, or 25 billion Fast tokens, at API rates. Anywhere below that, the API comes out cheaper. If you pick Vault, the real reasons are isolation and guaranteed capacity, not the price.
  3. Not on Bedrock yet. Amazon Bedrock still lists Embed 4 at $0.12 per 1M tokens, with no Embed 5 SKU. OpenRouter carries no Cohere embedding models at all.

Why the token price is the least important number here

Let me run the math on a realistic support setup, since that is the world I spend my working days in. Say you have 2,000 help center articles (about 1,500 tokens each) and 200,000 past tickets (about 600 tokens each), roughly 123 million tokens in total. Embedding all of it once with Pro costs about $14.76. With Fast, about $9.84. Fifty thousand customer questions a month at 30 tokens each is 1.5 million tokens, about 12 cents on Fast. In other words, the embedding bill is basically a rounding error.

What does not round away is the storage, which scales with both dimensions and precision. Cohere's own example: a 2048-dimension float32 vector is 8 KB, a 1024-dimension int8 vector is 1 KB, and a 256-dimension binary vector is 32 bytes.

Bar chart of raw vector storage for 100 million chunks: 819 GB at 2048 dimensions float32, 102 GB at 1024 dimensions int8, 3.2 GB at 256 dimensions binary, 256x smaller
Bar chart of raw vector storage for 100 million chunks: 819 GB at 2048 dimensions float32, 102 GB at 1024 dimensions int8, 3.2 GB at 256 dimensions binary, 256x smaller

At 100 million chunks, the default float32 output is about 819 GB of raw vectors. On Pinecone's Standard plan (or a hosted vector store) at $0.33 per GB a month, that is roughly $270 a month before index overhead, every month. The same chunks at 1024-dim int8 are about 102 GB, roughly $34 a month. Embedding those 100 million chunks once (at about 500 tokens each) costs about $6,000 on Pro, a one-time charge. Pick your dimensions and precision before you index, because if you change them later it means re-embedding everything from scratch.

Cohere's recommendation lines up with this: "For most deployments, we recommend 1,024-dimensional int8 vectors as the ideal performance-efficiency point," and int8 "retains near-full-precision retrieval quality." Binary is the smallest of them, it loses some accuracy, and it does a good job as a fast first pass ahead of a reranker.

Upgrading from Embed 4: what to plan for

If you are running Embed 4 today, the upgrade is more than just swapping the model name. Things to plan for:

  • Re-index everything. Cohere says Pro and Fast share a space with each other. It says nothing about Embed 4 vectors being comparable to Embed 5 vectors, so assume they are not and budget a full re-embed. With the prices above that is usually cheap in dollars, though it is expensive in terms of engineering time.
  • Watch the default dimension. Embed 4 defaulted to 1536, Embed 5 defaults to 2048. If your vector database index is fixed at 1536, pass output_dimension=1536 (Embed 5 supports it) or rebuild the index.
  • Check your cloud. If you call Cohere through Bedrock or Oracle OCI, Embed 5 is not there yet. The launch channels are the Cohere API, Model Vault, Microsoft Foundry, and SageMaker.
  • Batch jobs. Cohere says batch embedding is available for large-scale ingestion, but the docs table lists only the standard Embed endpoint for v5 models, so confirm Embed Jobs support for your model before you design a bulk pipeline around it.

Before blaming (or crediting) the model, check the rest of the pipeline first. This reply on an r/Rag thread about fine-tuning embeddings is the most practical advice on model swaps I have come across:

Reddit

"more than I expected. but only after chunking was already clean. dirty chunks make every embedding model look bad."

That lines up with my own experience. If you want a worked example, the guide to semantic search over Zendesk Guide walks through the chunking side. A better embedding model helps most when chunking is already sound and you have hybrid search and a reranker in place.

Where Embed 5 fits, and where it doesn't

Embed 5 is a great pick if you are a platform team building your own search or RAG stack, especially over long, visually rich, or financial documents. The 128K context means fewer awkward chunk splits and the fused text-plus-image input takes care of slides and scanned pages. On top of that, the Pro/Fast split gives you a clean dial for speed versus quality. Pair it with Parse 5 to turn PDFs into Markdown and Rerank to sort the matches, and you have Cohere's whole retrieval stack. Cohere also announced that its managed search platform, Compass Cloud, is now in private beta.

Hacker News

"Cohere seems to be doing a lot on the search side this year with their parsing model, Compass Cloud announcement, and now Embed 5... exciting stuff"

It is the wrong level of the stack if your actual goal is an everyday tool, like an AI knowledge base for your team or an AI that answers customer questions from your help center and past tickets. An embedding model hands you vectors. You still need chunking, a vector database, a reranker, a generation model like GPT-6.1 Sol, guardrails, and a way to plug answers into your helpdesk.

If you want the deeper trade-offs, start with RAG versus plain LLMs. For help centers specifically, there is a separate guide on RAG versus fine-tuning.

The hardest lesson I took from running AI on live support queues has nothing to do with retrieval quality, it is about what happens when retrieval comes back empty. I have watched paying customers' bots answer real customers with confident, made-up claims because the knowledge base had nothing relevant and the model filled the gap from its training data. A better embedding model makes that rarer, but it does not make it impossible. The fix lives in the layer above: a hard fallback when nothing relevant is found, and testing on real tickets before launch.

Try eesel if you want the answers, not the pipeline

If you are reading about embedding models because of a support problem, eesel is the shortcut. eesel is an AI teammate platform, and its AI helpdesk teammate is the whole retrieval stack, already assembled: it connects to your help center, docs, and past tickets, and drafts or sends replies inside your helpdesk, Slack, or a shareable link. No vectors to size, no index to rebuild when a new model ships.

It plugs into the helpdesks most support teams already run. On Zendesk it answers from macros and solved tickets; on Freshdesk and Gorgias it works the same queue your agents do.

The eesel Activity view filtered to a Zendesk instance, listing conversations the AI teammate resolved or left pending
The eesel Activity view filtered to a Zendesk instance, listing conversations the AI teammate resolved or left pending

Before it touches a live customer, eesel simulates the rollout on your historical tickets, so you see the replies it would have sent and the resolution rate up front. That is how I would want any automated ticket resolution to earn its way onto a live queue.

And if you came here because you like working from a terminal, the eesel CLI lets you run the same teammate from the command line: connect an integration with eesel integrations connect, edit its standing instructions, approve or deny pending actions, and read every run with eesel activity. Every command returns JSON and supports --dry-run, so scripts and coding agents like Claude Code can drive it too. There is a full walkthrough on managing agents from a terminal, and a shorter take on the CLI for customer support.

Try eesel free on your own tickets.

Frequently Asked Questions

What is Cohere Embed 5?
Cohere Embed 5 is Cohere's embedding model family, launched on September 30, 2026. It turns text, images, or text and images together into vectors for search and RAG. It ships in two tiers, Embed 5 Pro (embed-v5.0-pro) and Embed 5 Fast (embed-v5.0-fast), both with a 128K-token context window and 100+ languages.
How much does Cohere Embed 5 cost?
On the Cohere API, Embed 5 Pro costs $0.12 per 1M text tokens and Embed 5 Fast costs $0.08. Image input is $0.40 per 1M image tokens on both. Dedicated Model Vault instances start at $3.00 an hour ($2,000 a month). For the wider Cohere price list, see this Cohere pricing breakdown.
What is the difference between Embed 5 Pro and Embed 5 Fast?
Pro is the higher-quality model, meant for offline indexing and quality-critical retrieval. Fast costs a third less, runs about 2.4x the throughput, and scores close behind Pro on Cohere's benchmarks. They share one embedding space, so you can index documents with Pro and run live queries with Fast without rebuilding your vector store.
Is Cohere Embed 5 better than OpenAI embeddings?
On Cohere's own benchmarks, yes by a wide margin: ViDoRe V3 scores 85.8 for Embed 5 Pro versus 75.5 for OpenAI text-embedding-3-large. OpenAI's embeddings API is cheaper at the small end ($0.02 for 3-small), and the scores come from a metric Cohere introduced the same day, so test on your own data before switching.
Do I need to re-embed my data to move from Embed 4 to Embed 5?
Plan for it. Cohere says Pro and Fast share a space with each other, but it makes no claim that Embed 5 vectors are compatible with Embed 4 vectors, so a full re-index is the safe assumption. The default output size also moved from 1536 to 2048 dimensions, so set output_dimension explicitly if your index schema is fixed.
Is Cohere Embed 5 available on Amazon Bedrock?
Not at launch. Embed 5 is on the Cohere API, Model Vault, Microsoft Foundry (in preview, price not yet published), and Amazon SageMaker. Bedrock still lists Embed 4 at $0.12 per 1M tokens. If you are weighing cloud options, the Cohere alternatives roundup covers the field.

Share this article

Kira

Article by

Kira

Kira is a writer at eesel AI with a Computer Science background and over a year of hands-on experience evaluating AI-powered customer service tools. She focuses on breaking down how helpdesk platforms and AI agents actually work so that support teams can make better buying decisions.

Related Posts

All posts →
Illustration of two people reviewing a Cohere Embed 5 pricing dashboard with the Cohere logo
Trending

Cohere Embed 5 pricing: what Pro and Fast really cost in 2026

Cohere Embed 5 pricing is $0.12 per 1M tokens for Pro and $0.08 for Fast, with images at $0.40. Here is the full table, Model Vault math, and the storage bill nobody prices in.

Rama AdiRama AdiOct 1, 2026
Illustration of a pufferfish conductor routing a request across a school of models
Trending

Sakana Fugu Max: how it works, pricing, and benchmarks

Sakana Fugu Max is a cost-performance AI model that orchestrates a pool of other models behind one API. Here is how it works, what it costs, and who it is for.

KiraKiraSep 14, 2026
Cohere Parse 5 review: a document being scanned and split into tables, text and charts
Trending

Cohere Parse 5 review: is the $1.50 document parser worth it?

A hands-on Cohere Parse 5 review: what parse-v5.0 does, its ParseBench scores, the $1.50-per-1,000-pages pricing, the real limits, and who should use it.

Rama AdiRama AdiAug 30, 2026
Document pages being turned into structured tables, text blocks and charts by a parsing model
Trending

8 best Cohere Parse 5 alternatives for document parsing in 2026

The best Cohere Parse 5 alternatives for turning documents into clean text: LlamaParse, Mistral OCR, GPT-5.5, Gemini Flash, Textract and more, ranked and priced.

Kurnia KharismaKurnia KharismaAug 30, 2026
Cohere Parse 5 pricing breakdown illustration with the Cohere logo
Trending

Cohere Parse 5 pricing: what the $1.50 document parser really costs

Cohere Parse 5 costs $1.50 per 1,000 pages on the API, or a flat $2,500-$4,300/month on a dedicated instance. Here's the full breakdown and the break-even math.

Kurnia KharismaKurnia KharismaAug 30, 2026
Banner image for the Suno v6 review
Trending

Suno v6 review: I tested the new AI music models

A hands-on Suno v6 review: how v6, v6-wild, and v6-mini actually sound, the licensed-data pivot with Warner and BMG, what early testers say, pricing, and whether it beats v5.5.

KiraKiraSep 11, 2026
Illustrated hero banner for TypeSafe Jev, the ultrafast System One AI model, with a speed gauge
Trending

Is Jev really ultrafast? TypeSafe's System One model, tested

TypeSafe calls Jev an ultrafast System One model at 70-500ms a decision. Here is what the speed claim really means, where it holds up, and where it does not.

Rama AdiRama AdiSep 22, 2026
TypeSafe Jev pricing hero banner in rose and off-white, showing a low token cost per million
Trending

TypeSafe Jev pricing (2026): $0.042 per million tokens, output free

TypeSafe Jev pricing broken down: $0.042 per million input tokens, output free, no plan tiers yet, and what a System One model actually costs to run in production.

Kurnia KharismaKurnia KharismaSep 22, 2026
TypeSafe Jev hero banner in rose and off-white, illustrating a fast typed-decision model
Trending

TypeSafe Jev review: the 'System One' model that gives AI the properties of code

A hands-on TypeSafe Jev review: what the System One model actually does, whether the speed, price and 'can't hallucinate' claims hold, and where a typed-decision model fits real work.

Rama AdiRama AdiSep 21, 2026

Ready to hire your AI teammate?

Set up in minutes. No credit card required.

Get started free