Cohere has released Embed 5, a new embedding model family designed for enterprise search, RAG, and agentic retrieval. The release comes with two tiers: Embed 5 Pro for maximum retrieval quality and Embed 5 Fast for lower latency and cost. Both versions accept text, images, and fused text plus image inputs. They cover over 100 languages and process up to 128K tokens. The core design choice is that Pro and Fast share a single embedding space, allowing you to index with one and query with the other.
In this article
Both tiers are generally available on the Cohere API and Model Vault, Microsoft Foundry, and Amazon SageMaker. Private VPC or on-premises serving is possible via vLLM.
What Cohere Shipped
The API model IDs are embed-v5.0-pro and embed-v5.0-fast. Both output 2048, 1536, 1024, 768, 512, or 256 dimensions, with 2048 as the default. Embeddings return as float, int8, or binary. Pro costs $0.12 per 1M text tokens. Fast costs $0.08. Image inputs cost $0.40 per 1M tokens on both tiers.
Embed 5 can embed a page image directly. It can also fuse an image with its metadata into a single vector. This matters for scanned pages, slide decks, schematics, and charts, where text extraction loses information.
Pro and Fast: One Index, Two Query Paths
Cohere tested every corpus and query pairing across 40 development datasets. Normalised to Pro plus Pro at 100, a Pro index queried with Fast scored 98.4. An all-Fast setup scored 96.6. Cohere’s recommended pattern is to index with Pro and query with Fast. One constraint: both sides must use the same output dimension.
The split targets agentic workloads. An agent may issue dozens of searches per task, and query latency compounds. Cohere’s team reports Fast processed 377.3 documents per second versus 159.7 for Pro.
Benchmarks
On ViDoRe V3, Embed 5 Pro averages 85.8, an 8.8-point gain over Embed 4. Fast averages 84.5. Voyage 4 Large scores 83.7, Gemini Embedding 2 scores 83.2, and OpenAI text-embedding-3-large scores 75.5. On Cohere’s parsed-PDF suite, Pro leads at 84.8 against Voyage 4 Large at 83.6.
Finance is the strongest showing. Pro ranks first on FinanceBench (80.1), FinQA (90.0), and ViDoRe V3 Finance (85.0). Fast ranks second on all 3.
Multilingual results are mixed. Pro leads the European-language average at 77. However, Gemini Embedding 2 beats Pro on 9 of 10 further languages in Cohere’s own results table. Those include Japanese, Arabic, Hindi, and Telugu.
One important thing to note. Most numbers use RCP-nDCG@10, a new Cohere metric. It reorders a fixed candidate set, so it measures reranking quality more than first-stage retrieval. Cohere published the evaluation code, but independent replication is still pending.
Storage Costs at Scale
Embed 5 uses Matryoshka representation learning plus lower-precision outputs. A 2048-dim float32 vector needs 8 KB. A 1024-dim int8 vector needs 1 KB. A 256-dim binary vector needs 32 bytes. Across 100M chunks, raw storage drops from about 819 GB to 3.2 GB. Cohere recommends 1024-dim int8 as the default, citing near-full-precision quality.
How Embed 5 Works
1. What an embedding model actually does
Text goes in, a vector of numbers comes out. Search then ranks documents by how close their vectors sit to the query vector.
Query: What happened to net interest margin last quarter?
Documents ranked by cosine similarity:
- Net interest margin narrowed 12 bps to 2.61% as deposit costs rose.
- Torque the mounting bolts to 45 Nm in a star pattern before refitting the cover.
- Employees accrue 1.5 days of paid leave for each month of service.
Documents come from Cohere’s launch code sample. Vector bars and similarity scores here are illustrative, not real model output.
2. Two tiers, one embedding space
Pro and Fast write vectors into the same space. Pick which model indexes the corpus and which one embeds the query. Quality is normalised so Pro plus Pro equals 100.
Cohere’s recommended pattern is to index with Pro and query with Fast.
Document throughput (docs per second, Cohere-measured):
- Embed 5 Fast: 377.3
- Embed 5 Pro: 159.7
Price per 1M text tokens: Pro $0.12, Fast $0.08. Images cost $0.40 per 1M tokens on both tiers.
3. Shrink the vector, not the quality
Matryoshka training lets you truncate dimensions. Lower precision outputs cut bytes further. Cohere recommends 1024-dim int8 as the default sweet spot. Binary trades some accuracy and suits a first pass before reranking.
4. How it scores against rivals
Cohere-reported averages using its new RCP-nDCG@10 metric. Treat these as vendor numbers until independent runs land.
Source: Cohere Embed 5 launch post, Sep 30 2026. RCP-nDCG@10 reorders a fixed candidate set, so it reflects reranking quality more than first-stage recall.
5. How much fits in one input
A longer context means fewer chunks for long filings and manuals. Limits taken from each vendor’s model docs. 128K is shown as 131,072 tokens for bar scale.
- Cohere Embed 5 Pro / Fast: 128K tokens
- Voyage 4 Large: 32,000 tokens
- Gemini Embedding 2: 8,192 tokens
- OpenAI text-embedding-3-large: 8,191 tokens
6. Is it deployable? Yes. Pick a route.
Both tiers are generally available.
- Cohere API: Managed endpoint
- Model Vault: Dedicated inference
- Microsoft Foundry: Azure catalog
- Amazon SageMaker: AWS Marketplace
- Private VPC / on-prem: vLLM




