Eight months ago I picked our vector database for a RAG pipeline based on a benchmark with one million vectors. Our production index crossed 40 million by June. Recall dropped, p99 latency crept from 80ms to just under 400ms, and I spent a week convinced it was a filtering bug before I accepted the real answer: the benchmark I’d trusted never tested anything close to our actual scale, and I didn’t think to ask.
I wasn’t alone in that. On September 1, 2026, Qdrant published a blog post titled “Enough with the Bad Benchmarks,” and it’s the first time I’ve seen a vector database vendor say the quiet part out loud: most public vector search benchmarks are synthetic, undersized, or run against gated services nobody can independently verify. Their fix is Qdrant-FineWeb-10B — 10.07 billion documents, 24.47 TB of vector data, 28.66 TB of source text and metadata, built from a slice of Hugging Face’s FineWeb corpus and embedded with gte-multilingual-base.
Why the ground truth actually matters
The part that made me sit up wasn’t the document count. Anyone can generate 10 billion random vectors. The expensive, credible part is the ground truth: Qdrant computed exact top-1,000 nearest neighbors for 100,000 dense, sparse, and filtered queries by brute-force comparison against the full 10-billion-vector corpus — over one quadrillion distance calculations, done on GPU infrastructure. That’s the number that matters, because approximate nearest neighbor indexes (HNSW, IVF, whatever your database uses under the hood) are only as good as what you’re measuring recall against. If your “ground truth” is itself approximate, your recall numbers are fiction wearing a lab coat.
They also open-sourced the tool that did this work, called Supernova, split into four pieces: nova-embed for generating embeddings, nova-bf for GPU-native brute-force ground truth, nova-load for getting data into a target database, and nova-dist for distributing the whole pipeline across a cluster via SkyPilot. If you’ve ever tried to build your own eval harness for a vector database migration, you know how much of this work is unglamorous plumbing that nobody wants to write twice. Having it as a released tool, not just a paper, is the actually useful part.
What I did with it
I pulled a small slice — not the full 10B, my laptop and my patience both have limits — to sanity-check the filtered-query pattern that’s closest to what bit us in production: fetch top-k nearest neighbors with a metadata filter (tenant ID plus a date range) applied at query time, not post-filtered.
from qdrant_client import QdrantClient
from qdrant_client.models import Filter, FieldCondition, Range
client = QdrantClient(url="http://localhost:6333")
# Filtered ANN query pattern, same shape as our prod hot path
result = client.query_points(
collection_name="fineweb_slice",
query=query_vector,
query_filter=Filter(
must=[
FieldCondition(key="tenant_id", match={"value": "t_4471"}),
FieldCondition(key="crawl_date", range=Range(gte=1735689600)),
]
),
limit=20,
with_payload=False,
)
At 1M vectors, this pattern returns in single-digit milliseconds and recall against brute-force ground truth sits comfortably above 98%. At 50M vectors with the same filter selectivity, recall on our HNSW index dropped to around 89% before we retuned ef_search upward — and that retune cost us roughly 35% more query latency to buy the recall back. None of that shows up if your only benchmark data point is “works great at 1M vectors,” which is exactly the size most public benchmarks, and most of our own pre-launch testing, actually used.
The uncomfortable part: what this means for how we test
The honest version of this story is that we didn’t skip benchmarking. We benchmarked carefully, at a scale that felt reasonable for a six-month-old product, and the benchmark told us the truth about that scale. It just didn’t tell us about the scale we’d hit eight months later, because no cheap benchmark can simulate 40 million vectors on a laptop, and nobody budgets GPU-hours for “let’s brute-force ground truth at production scale” during a sprint planning meeting.
Qdrant’s quote in the release nails the actual bar: “Engineers within the world’s leading search teams demand reproducibility, 95%+ recall, large retrieval depth, high throughput, and sub 100-ms p99 latency.” Every one of those is a scale-dependent property. Recall at 1M tells you nothing reliable about recall at 100M, because the neighbor density and index approximation error both shift nonlinearly with corpus size — HNSW’s approximation gets worse in ways that aren’t a clean function of size, and each database’s default tuning was almost certainly validated at whatever scale the vendor benchmarks internally, which for most public marketing benchmarks is not billions.
What I’m changing on our side: for any vector index projected to grow past 10x its launch size within a year, we now budget an actual scale re-benchmark before launch, not just at the point where p99 alerts start firing. That means picking a realistic subset of FineWeb-10B or Coyo-Vector-Embeddings (their 15.4M-vector multimodal set using Qwen3-VL-Embedding-2B, closer to what we’d need for an image-heavy pipeline) sized to our year-one growth target, not our launch-day size.
The takeaway that isn’t a bow
A benchmark answers the question you asked it, at the scale you asked it at. Ours answered “does this work at 1M vectors” perfectly. It was never asked “does this work at 40M,” and a benchmark that was never asked a question will never volunteer the answer. If you’re picking vector infrastructure right now, go pull Qdrant-FineWeb-10B (released September 1, 2026) or at least read past the headline number — the ground-truth methodology is the part worth stealing for your own eval pipeline, database-agnostic, whether you end up on Qdrant or not.