
Cohere Embed 5 pricing at a glance
Cohere launched the Embed 5 family on September 30, 2026, in two tiers, and both of them share the same 128K context and 100+ languages, plus inputs that can be text, image, or the fused text-plus-image kind. I went through every published price and put them into one table, taken straight from Cohere's pricing page, the launch post, and the cloud marketplace listings:
| Option | Price | Billed by | Best for |
|---|---|---|---|
| Trial key | Free, 1,000 calls/month | n/a (no production use) | Testing |
| Embed 5 Pro, text | $0.12 / 1M tokens | Input tokens | Indexing documents |
| Embed 5 Fast, text | $0.08 / 1M tokens | Input tokens | Queries, agent loops |
| Embed 5 Pro or Fast, images | $0.40 / 1M tokens | Image tokens | Page images, slides |
| Model Vault Small (Pro or Fast) | $3.00/hr or $2,000/month | Dedicated instance | Steady high volume |
| Model Vault Medium (Pro or Fast) | $5.00/hr or $3,250/month | Dedicated instance | Very high volume |
| Amazon SageMaker | $2.39 to $8.48/host/hr software fee + instance | Hour | Teams already on AWS |
| Microsoft Foundry | Listed in preview, no price yet | n/a | Teams on Azure |

The same tab also lists Parse 5 at $1.50 per 1,000 pages and Rerank 4 from $2.00 per 1,000 searches, and together with Embed 5 those make up the full retrieval stack Cohere sells. Reading the cards, there were two things that jumped out for me. First one, Model Vault charges the same rate for Pro and Fast, which means on a dedicated instance the whole "Fast is cheaper" logic just disappears, and the only reason to pick Fast there is the speed. Second, Embed 4 is gone from the pricing page entirely, even though the models docs still list it with no deprecation notice. If you're on Embed 4 today nothing breaks, but the signal about where Cohere wants you to be is pretty clear. (My Cohere AI review covers the rest of the lineup.)
I build integrations at eesel, and the part of eesel that turns a help center and a few years of tickets into something an AI can search is exactly this layer. So when I read embedding price cards, I read them the way a plumber reads a quote for pipes. The per-unit price matters, sure, but what you are really pricing is everything that the pipes connect to.
What changed from Embed 4
The short version of it: Cohere held the line on price, and then moved pretty much everything else. Here's the side-by-side, using Embed 4's rates from Azure's Cohere pricing and Cohere's August pricing snapshot:
| Embed 4 | Embed 5 Pro | Embed 5 Fast | |
|---|---|---|---|
| Text, per 1M tokens | $0.12 | $0.12 | $0.08 |
| Images, per 1M tokens | $0.47 | $0.40 | $0.40 |
| Model Vault Small | $4.00/hr, $2,500/mo | $3.00/hr, $2,000/mo | $3.00/hr, $2,000/mo |
| Default dimensions | 1536 | 2048 | 2048 |
| ViDoRe V3 score (Cohere) | 77.0 | 85.8 | 84.5 |
So you get an 8.8-point quality jump at the same text price, images 15% cheaper, and a dedicated instance $500 a month cheaper. Fast, on the other hand, didn't exist at all before.
The row I would circle is the default dimensions one. The price per token is the same, but each vector comes out a third bigger by default. If you swap models and keep the default settings, your embedding bill stays flat and your vector database bill goes up. I'll get to the fix further below, since it's an easy one and it does save real money.
What counts as a billable token
Cohere's billing docs say embed is "priced based on the number of tokens embedded", and tokens Cohere adds behind the scenes aren't charged. You only pay for the input. There is no output-token charge at all, since what an embedding call returns is a list of numbers and not text.
There are a few details worth knowing that shape the real bill:
- Batching doesn't change the price. Each embed call takes up to 96 texts or inputs, which is good for throughput but the cost per token stays the same.
- Images are priced per token, but the conversion isn't published. The docs explain resizing (images over 2,458,624 pixels get scaled down) yet give no tokens-per-image formula. So it's worth running a small sample first, before you put a budget on a large PDF archive.
- There's no stated batch discount. The launch post says batch embedding is available, but the Embed Jobs guide still says it only works with v3.0 models. Don't plan around a 50% batch rate like the one you'd get from Google.
- Trial keys are narrow. Text embedding runs at 2,000 inputs a minute on both trial and production, but image embedding is 5 inputs a minute on trial vs 400 on production.
Everything else after that is plain arithmetic, tokens in times the rate. It's the same billing model as the OpenAI API and most other hosted embedding services.
The index with Pro, query with Fast trick
For me this is the most useful thing in the whole launch, and it's really a pricing feature that is dressed up as a technical one. Pro and Fast share one embedding space, so a query that Fast embedded can search an index that Pro built, without you rebuilding anything.
"For many customers, we recommend the following deployment pattern: index with Pro, query with Fast."
Cohere ran a test on every pairing across 40 development datasets. With an all-Pro setup scored as 100, a Pro index queried by Fast came in at 98.4, a Fast index queried by Pro at 97.3, and all-Fast at 96.6. There is one catch hiding in the footnotes, which is that both sides have to use the same output dimension.

So why does it matter for the cost? It's because in most search and RAG systems (here's a support RAG pipeline if you want the full shape), indexing happens once (plus updates), while queries happen forever. That way you pay the Pro rate on the part you do rarely, and the Fast rate on the part that you do every second. Fast also handled 377.3 documents a second vs Pro's 159.7 in Cohere's throughput test, so queries get quicker as well as cheaper.

Here's where I would push back a bit on the common read, though, because for most teams the query side is tiny. A support bot handling 30,000 questions a month at about 100 tokens each sends 3 million query tokens: that's $0.36 on Pro or $0.24 on Fast. The "Fast saves money" headline is real at agent-loop scale, where a model fires dozens of searches per task. For a normal help center, pick Fast for the latency and then stop worrying about the cost of it.
The cost the pricing page doesn't show: vector storage
Every vector you create has to live somewhere, usually a vector database, and that somewhere bills you monthly. Out of a whole embedding budget, this is the part I see underestimated the most, because the embedding API is a one-time cost per document while the database is rent.
Embed 5 lets you choose output size (256, 512, 768, 1024, 1536, or 2048 dimensions) and format (float, int8, binary, and a few variants). Cohere's own example: 100 million chunks at the 2048-dimension float default take 819 GB, while 256-dimension binary takes 3.2 GB, which is 256x smaller.

I priced those three sizes on Pinecone's serverless plan at $0.33 per GB per month (raw vector bytes only, before index overhead and metadata). That puts the gap at $270 a month vs $34 vs about $1. Over two years, the default float setting costs roughly $6,500 in storage alone, more than the $6,000 it takes to embed 50 billion tokens with Pro in the first place.
What you lose in quality by shrinking is smaller than you'd guess. Cohere's chart plots retrieval quality against relative storage cost, and the int8 lines sit almost on top of the float lines. Cohere's own recommendation is 1024-dimension int8, and it's also the setting I would start with.

Other databases price things differently, and the shape of the bill changes along with them. Weaviate Cloud charges per million vector dimensions stored (from $0.00465 on Flex), so 2048 dimensions literally costs double 1024. Pinecone also charges $16 to $18 per million read units on Standard, with a $50 monthly minimum, and for small indexes that matters more than the storage does. I compared the database options in more depth in my Weaviate pricing breakdown and the Weaviate alternatives roundup.
If you'd rather not run a database at all, OpenAI's hosted option is covered in my vector stores guide, though it ties you to OpenAI's own embedding models.
Estimate your Embed 5 bill
Plug in your own corpus and your own traffic here. The calculator uses the published API rates and Pinecone's $0.33/GB storage rate, so the storage line is best treated as a floor.
Try switching from 1024-dimension int8 to 2048-dimension float, which are Embed 5's defaults. On a million chunks the indexing bill doesn't move at all, while the storage line jumps 8x.
API vs Model Vault vs cloud marketplaces
Model Vault swaps the per-token meter for a flat rent on a dedicated, single-tenant instance. The break-even math on it is simple enough:
| Instance | Monthly | Breaks even on Pro | Breaks even on Fast |
|---|---|---|---|
| Small | $2,000 | ~16.7B tokens/month | ~25B tokens/month |
| Medium | $3,250 | ~27.1B tokens/month | ~40.6B tokens/month |
Sixteen billion tokens a month is about 550 million a day. Very few teams embed that much, so for almost everyone the API stays the cheaper option. Cohere also doesn't publish how many tokens a Small instance can push per month, so check the throughput with sales before assuming one instance will cover your peak.
People buy Model Vault anyway, and the reason is control more than price. There's no shared tenancy and throughput is guaranteed, which also gives you a cleaner story for security reviews. That fits Cohere's whole pitch around enterprise deployment. If you need it to sit fully inside your own network, both models can be self-hosted on vLLM through a private deployment, with the price set by sales.
Then the marketplaces are a third path. On Amazon SageMaker, Embed 5 bills as an hourly software fee per host ($2.39 on an ml.g5.xlarge, up to $8.48 on the largest Pro instance) plus the SageMaker instance itself. Microsoft Foundry lists both models in preview without a price yet. Amazon Bedrock and Oracle OCI still only carry Embed 4 as of October 1, at $0.12 per 1M tokens on Bedrock. If your company has committed cloud spend, the marketplace route can still make sense even when the raw rate is worse.
Worked examples: what a real team pays
Below are three scenarios, using the published rates and 1024-dimension int8 storage unless I note otherwise.
A support team's help center and ticket history. 2,000 help articles at about 800 tokens and 50,000 past tickets at about 400 tokens comes to 21.6 million tokens. Indexing with Pro costs $2.59, once. Queries at 30,000 a month on Fast cost $0.24 a month. The vectors fit inside Pinecone's free 2 GB tier. The embedding bill in this case is a rounding error. Where the real cost sits is the engineering to build chunking and syncing, plus permissions and the LLM layer on top, which is the same trade-off I walk through in semantic search over Zendesk Guide.
A document archive for an internal search tool. 20 million chunks at 500 tokens is 10 billion tokens. Pro indexing costs $1,200; Fast would be $800. Storage at 2048-dimension float is about 164 GB, or $54 a month on Pinecone. At 1024 int8 it's about 20 GB, or $6.76 a month. When the next model ships, a re-index costs the full $1,200 again, and if the PDFs need parsing first, Parse 5 adds its own per-page bill.
A high-volume agent that searches on every step. 500 million query tokens a day is 15 billion a month. On Fast that's $1,200 a month; on Pro, $1,800. Neither of these crosses the $2,000 Model Vault line, so the API still wins unless dedicated capacity is something you need.
The lesson across all three: the embedding API is almost never the expensive part. Storage and re-indexing are, and so are the people maintaining the pipeline.
How Embed 5 pricing compares
I lined Embed 5 up against the other flagship embedding models, each price taken from the vendor's own pricing page. The scores are Cohere's ViDoRe V3 numbers from its launch post, so read them as Cohere's own test and not as an independent one.

| Model | Text / 1M tokens | Batch | Free tier | ViDoRe V3 (Cohere) |
|---|---|---|---|---|
| Cohere Embed 5 Pro | $0.12 | Not stated | 1,000 calls/mo trial | 85.8 |
| Cohere Embed 5 Fast | $0.08 | Not stated | 1,000 calls/mo trial | 84.5 |
| Voyage 4 Large | $0.12 | 33% off | 200M tokens | 83.7 |
| Gemini Embedding 2 | $0.20 | $0.10 | Yes | 83.2 |
| OpenAI text-embedding-3-large | $0.13 | n/a | None | 75.5 |
| OpenAI text-embedding-3-small | $0.02 | n/a | None | n/a |
| Jina Embeddings v5 | $0.045 to $0.05 | n/a | 10M tokens, non-commercial | 74.5 (v5 Small) |
| Mistral embed | $0.10 | n/a | n/a | n/a |
Embed 5 Fast is the cheapest model in the top-scoring group, and Pro matches Voyage 4 Large on price while edging slightly ahead of it on Cohere's test. Gemini Embedding 2 costs 2.5x what Fast does at full rate, though its $0.10 batch price narrows that a lot for offline indexing (full rates in my Gemini pricing guide). The OpenAI embeddings API is a cent more than Pro for a much lower score on this benchmark.

There are two honest caveats to flag here. If pure price per token is your only goal, OpenAI's 3-small and Voyage 4 Lite both sit at $0.02, a quarter of Fast. And Cohere's own multilingual table shows Gemini Embedding 2 ahead of Pro on 9 of the 10 Asian and Middle Eastern languages it tested, with Telugu the widest gap (91 vs 80). If your corpus is mostly in Hindi, Japanese, or Arabic, it's worth testing Gemini before you commit. Mistral's embed model at $0.10 is another option for European data; see my Mistral AI pricing notes. For a broader list, my Cohere AI alternatives post compares the vendors beyond embeddings.
What developers are saying
Embed 5 is a day old as I write this, so there aren't real user reviews of it yet. What does exist is a steady track record for Cohere's embeddings in general, and the theme that keeps coming up is reliability:
"It has the most crisp, steady P50 of any external service I've used in a long time."
The pushback tends to be about cost at volume and about depending on a closed API, which is true for every hosted embedding model:
"The API can also become expensive when you start using it more frequently."
"Voyage for embeddings + rerank is totally defensible - Cohere is popular partly because it's been the default in a lot of examples, not because it's always better."
I think that last one is fair, and the Embed 5 numbers back up both sides of it: Pro wins Cohere's own benchmark, but by 2.1 points over Voyage 4 Large at the same price. If you already run Voyage well, there's no urgent reason for migrating. If you're on Embed 4, moving to Embed 5 costs nothing extra per token and is close to a pure win. The LangChain Cohere integration makes the swap a one-line model change, though you'll still need to re-embed the corpus.
Which Embed 5 pricing path fits you
Here are my recommendations, by situation:
- You're prototyping. Use the trial key, then a pay-as-you-go production key. Don't think about Model Vault yet.
- You're building search or RAG over a fixed corpus. Index with Pro, query with Fast, store at 1024-dimension int8. This is the cheapest high-quality setup Cohere offers.
- You're running agents that search constantly. Fast on both sides if you need the 2.4x throughput; Pro index plus Fast queries if quality matters more.
- You embed tens of billions of tokens a month or need isolation. Price Model Vault, and ask sales for throughput per instance before you sign. On AWS, compare it with SageMaker and the Bedrock agent pricing you may already pay.
- You're mostly embedding Asian-language content. Benchmark Gemini Embedding 2 against Pro on your own data first.
- You just want a support bot that answers from your docs. Skip the embedding decision entirely and look at an AI knowledge base or helpdesk teammate that handles retrieval for you.
Try eesel when the goal is support answers
Most people who are pricing embedding models aren't building a search engine just for its own sake. They're often choosing between RAG and fine-tuning for a help center. What they want in the end is a bot that answers customers from a help center and past tickets. That's the point where I would stop and ask if you need to own the stack at all.
At eesel I've watched teams make this call from both sides. One Dutch web-hosting company on Zendesk, spending about $1,700 a month, left eesel to build their own AI in-house. The read from the team afterwards had nothing to do with model costs:
"A lot of their friction came from setup complexity (handovers, Zendesk integration, business hours logic) and a lack of clear guidance on how to configure things correctly."
That's the part embedding pricing pages never show. The vectors are cheap. It's the handoffs and integrations, plus the rules around them, that are the real work.

eesel is an AI helpdesk teammate that joins your existing queue in Zendesk, Freshdesk, and other helpdesks, learns from your knowledge base and past tickets, and lets you replay real tickets through it before it talks to a customer. There's no embedding model to pick and no dimension count to tune, and you don't get a vector database bill either.
Pricing is fixed credit plans starting at $299 for 500 credits a month, where a ticket or chat handled is one credit.
If you're a developer, going this way doesn't mean you lose control, which I cover in more depth in my customer support agent API guide. The eesel CLI lets you connect integrations, edit the teammate's instructions, approve or deny pending actions, and read its activity log from a terminal, with JSON output and a --dry-run flag for scripting. Coding agents like Claude Code are able to drive it over MCP too.
I've written more about that workflow in managing agents from terminal and the CLI for support walkthrough.
Try eesel free and see what your help center looks like as a working support teammate, without writing a single embedding call.
Frequently Asked Questions
How much does Cohere Embed 5 cost?
What is the difference between Embed 5 Pro and Embed 5 Fast pricing?
Is Cohere Embed 5 more expensive than Embed 4?
Is there a free tier for Cohere Embed 5?
When does Cohere Model Vault become cheaper than the Embed 5 API?
How does Cohere Embed 5 pricing compare to OpenAI embeddings?
Do I need Cohere Embed 5 to build an AI support agent?

Article by
Rama Adi
Rama is a software engineer at eesel AI with two years of experience writing about B2B SaaS, AI tools, and customer support technology. Based in Bali, Indonesia, he brings a developer's perspective to product comparisons — cutting through marketing copy to what the integrations and APIs actually do.








