Cohere's Embed 5 splits indexing and querying into separate models without rebuilding vectors
Cohere released Embed 5 on Wednesday, allowing teams to index data with the Pro model and query using the faster, cheaper Fast variant from the same embedding space. The company's tests show minimal retrieval quality loss when mixing the two models.

Cohere unveiled Embed 5 this week, introducing a two-tier embedding strategy that lets organizations index their data with Embed 5 Pro while performing queries through Embed 5 Fast—all without needing to regenerate embeddings. This approach particularly benefits RAG and agent systems, where repeated searches over the same indexed material can accumulate latency costs.
The company's recommendation pairs Pro for the indexing phase with Fast for the query phase, a split that makes economic and performance sense in workloads where documents are ingested infrequently but searched constantly. Cohere's evaluation across 40 datasets spanning text, images, fused documents, and parsed documents found that Fast queries against a Pro index achieved a score of 98.4 relative to a Pro-to-Pro baseline of 100. When Fast handled both indexing and queries, the score dropped to 96.6. According to Cohere, no individual dataset showed substantial degradation when Pro and Fast worked together.
One embedding space, two models
The technical foundation enabling this flexibility is a shared embedding space between Pro and Fast. Teams can transition between the two models without re-embedding their corpus since both generate compatible vectors at identical dimensions. Cohere indicates that Matryoshka truncation and int8 quantization remain compatible when mixing the two variants.
Because Pro and Fast share an embedding space, teams can switch between them without re-embedding the corpus.
Cohere
Pricing favors the Fast model at $0.08 per million tokens against Pro's $0.12 per million tokens, with Fast delivering roughly 2.4 times the document throughput in Cohere's benchmarks. For RAG systems where document ingestion happens less frequently than searches, Pro can process incoming documents while Fast manages the substantially heavier query load.
Shrinking vectors with Matryoshka
Both models support six vector dimensions ranging from 256 to 2,048, available in float32, int8, and binary formats. Storage implications become pronounced at scale, particularly for organizations reconsidering their vector infrastructure location.
Cohere estimates a 2,048-dimensional float32 vector at 8 KB, translating to roughly 819 GB for 100 million chunks. Reducing to a 1,024-dimensional int8 vector cuts this to approximately 102 GB, while a 256-dimensional binary vector compresses the same corpus to roughly 3.2 GB.
For typical deployments, Cohere suggests 1,024-dimensional int8 vectors, which reduce memory and storage demands while preserving near-full-precision retrieval performance. Binary representations provide more aggressive compression at the cost of some accuracy and work better as an initial retrieval stage before higher-precision reranking.
Both models support six vector dimensions from 256 to 2,048, with float32, int8 and binary formats.
Cohere
Retrieval beyond plain text
Embed 5 handles text, images, and combined text-image inputs across more than 100 languages, with a 128K-token context window. The model can embed page images directly or fuse image and text inputs into a single vector.
On Cohere's five-dataset fused text-image evaluation, Pro averaged 82.3 compared with Fast at 81.2 and Google's Gemini Embedding 2 at 61.3. For parsed-PDF evaluation, Pro scored 84.8, trailed by Voyage 4 Large at 83.6, Fast at 83.4, and Gemini Embedding 2 at 80.8.
On ViDoRe V3, where Cohere used parsed text outputs curated by the benchmark authors rather than page images, Pro averaged 85.8 and Fast 84.5, against 83.7 for Voyage 4 Large, 83.2 for Gemini Embedding 2, and 77 for Cohere's previous Embed 4.
Cohere's multilingual performance shows less clear dominance, a notable point given the company's recent focus on machine translation. Pro led its five-language European average with 77, compared with 76 for Voyage 4 Large and 73 for Gemini Embedding 2, but trailed Gemini Embedding 2 on nine individual tests.
Reading the benchmark fine print
Embed 5 marks Cohere's first model family evaluated using RCP-nDCG@10, a metric employing query-specific relevance criteria instead of fixed relevance labels.
Cohere contends this approach identifies relevant results that original benchmark labels might overlook, though RCP-nDCG@10 measures reranking performance over a fixed candidate set rather than first-stage retrieval from a complete corpus.
First-stage retrieval gets evaluated separately using standard nDCG and Recall metrics, while fused text-image, page-image, and cross-model tests employ standard nDCG@10. These differences mean reported scores lack direct comparability across evaluations.
Separating indexing from serving
The more significant shift in Embed 5 is decoupling indexing from serving as independent infrastructure choices. A single corpus can be indexed for retrieval quality while the query path optimizes for throughput and latency, without maintaining duplicate data representations.
Cohere's 98.4 score indicates the Pro-to-Fast configuration sacrifices relatively modest retrieval quality in its tests, though this figure averages across Cohere's evaluation suite. Real-world RAG and agent systems should still benchmark Pro-to-Fast against Pro-to-Pro using their own data and query patterns, especially when retrieval mistakes can propagate through multiple agent workflow stages.
Embed 5 Pro and Fast are accessible through Cohere's API and Model Vault, Microsoft Foundry, and Amazon SageMaker, with private VPC and on-premises options available via vLLM.
Production RAG and agent systems will still need to benchmark Pro-to-Fast against Pro-to-Pro on their own corpus and query distribution, particularly when retrieval errors can carry through multiple steps of an agent workflow.
Cohere