Every RAG system I’ve shipped has the same uncomfortable constraint baked in on day one: whatever embedding model you index with, you query with. Forever, or until you eat the cost of re-embedding your entire corpus. Cohere’s Embed 5, released September 30, breaks that constraint on purpose. embed-v5.0-pro and embed-v5.0-fast share a vector space. You can index with Pro and query with Fast — or the reverse — without re-indexing anything.
That’s not a benchmark flex. It’s a different default architecture for anyone running retrieval at a scale where query volume dwarfs document volume, which is most production RAG.
The numbers, not the marketing
Across 40 datasets spanning text, images, fused documents, and parsed documents, Cohere measured four combinations against a Pro-index/Pro-query baseline of 100:
- Pro index, Fast query: 98.4
- Fast index, Fast query: 96.6
- Cohere’s own write-up: cross-model mismatch costs “1.6% and 2.7% losses for Fast and Pro queries, respectively,” and that number holds up across Matryoshka truncation and int8 quantization — so you can also shrink vectors without the cross-model penalty compounding on top.
On ViDoRe V3 (built around visually dense enterprise documents — the stuff with tables, scanned PDFs, mixed layouts), Embed 5 Pro scores 85.8, Fast scores 84.5, both ahead of Voyage 4 Large at 83.7. Throughput: Fast does 377.3 documents/second against Pro’s 159.7 — 2.4x.
Pricing is where the architecture decision actually gets made:
| Model | Text | Image |
|---|---|---|
| Pro | $0.12 / M tokens | $0.40 / M tokens |
| Fast | $0.08 / M tokens | $0.40 / M tokens |
Image pricing is identical across tiers. Text is where the split pays off, and text is where query volume actually lives for most systems.
Why this changes where you spend your embedding budget
Indexing and querying have completely different cost profiles in a real system, and most teams accidentally optimize for the wrong one because they only ever had one model to tune.
Indexing: happens once per document.
volume: bounded by corpus size
latency tolerance: high (batch job, overnight, whatever)
quality tolerance: zero — a bad embedding here is permanent
until you re-index
Querying: happens once per user request.
volume: unbounded, scales with traffic, not corpus size
latency tolerance: low (user is waiting)
quality tolerance: some — reranking downstream can absorb
a few points of recall loss
Before Embed 5, you picked one model and both profiles inherited its cost and latency characteristics whether they needed to or not. A support-ticket search system with 50,000 documents and 2 million queries a month was paying Pro-tier latency and Pro-tier price on every single one of those 2 million queries, because the document count never justified switching models and re-indexing.
With a shared vector space, the indexing side can afford to be expensive and careful — it’s a one-time cost amortized over the corpus’s lifetime — while the query side gets tuned for what it actually is: a high-volume, latency-sensitive, moderately-quality-tolerant path.
# index once, with the higher-fidelity model
doc_embeddings = co.embed(
model="embed-v5.0-pro",
input_type="search_document",
texts=documents,
output_dimension=1024,
embedding_types=["float"],
).embeddings.float_
# query with the cheap, fast model against the same index
query_embedding = co.embed(
model="embed-v5.0-fast",
input_type="search_query",
texts=[user_query],
output_dimension=1024,
embedding_types=["float"],
).embeddings.float_
Nothing else in the retrieval pipeline changes. Same vector store, same ANN index, same distance metric. The model choice moved from an architecture decision locked in at index time to a runtime parameter you can flip per request.
Where the 1.6-2.7% actually matters
That loss number isn’t free to ignore just because it’s small. Run the math against your own system before you flip the switch:
- High query volume, forgiving task (FAQ search, internal doc lookup, anything with a human skimming top-10 results): take the Fast-query discount without hesitation. A 1.6% relative drop in a top-10 retrieval set is noise compared to what your reranker or your UI’s “did you mean” already absorbs.
- Low query volume, precision-critical task (legal document retrieval feeding a generation step with no human review, compliance search where a missed document is a real liability): stay on Pro-to-Pro. The throughput gain buys you nothing if your query volume was never the bottleneck, and the quality floor matters more than the latency ceiling.
- The actual sweet spot — and this is the case most RAG systems in production fall into — is Pro-indexed documents queried with Fast, feeding a reranker. The reranker was already there to clean up retrieval noise; a 1.6% quality dip from the embedding stage is exactly the kind of error a cross-encoder reranker was built to catch. If you’ve already got a rerank step, as The New Stack’s coverage of the retrieval funnel pattern argued this same week, you’ve got slack in the pipeline specifically for this trade.
What I’d check before switching
I don’t trust a 40-dataset aggregate number to predict what happens on my corpus specifically — aggregate benchmarks smooth over the cases where a domain’s vocabulary or document structure happens to be exactly where a smaller model falls apart. Before moving a production query path from Pro to Fast:
- Pull your actual query logs, not a synthetic eval set, and run both models against your real index.
- Measure recall@10 and recall@20 specifically on the queries your support team has flagged as “search didn’t find the obvious answer” — that’s your adversarial set, and it’s worth more than any public benchmark.
- Check whether your downstream reranker or generation step is already correcting for retrieval imperfection. If it is, the 1.6-2.7% is close to free. If your pipeline has no reranker and the top-1 result goes straight into a prompt, that gap is exactly where a wrong answer gets generated with full confidence.
The headline here isn’t “Cohere’s new model is faster.” It’s that embedding model choice stopped being a single irreversible decision baked into your index, and started being a parameter you tune per code path based on what that path actually needs. That’s a more useful shape for a production system than another half-point of benchmark score.