JOURNAL
950 Agents, 21 Hours, 210M Tokens: What Anthropic’s CRISPR Search Actually Proves About Fan-Out
Anthropic's life-sciences lab ran 950 agents in parallel to surface a possible new CRISPR-like enzyme system. The funnel math is a genuinely useful fan-out pattern — but the scientist pushback is the part worth reading closely.

On this page
950 agents. 21 hours. 210 million tokens. That’s the detail everyone skips past on the way to the CRISPR headline, and it’s the only part of this story I actually trust without a biology degree.
Anthropic’s life-sciences team reported, via TechCrunch on September 23rd, that their agent fleet surfaced something they’re calling ART — array-associated reverse transcriptases, structurally reminiscent of CRISPR. The funnel: 200,000 candidate reverse transcriptases, filtered down to 3,500, filtered down to 20 systems, one of which is the finding people are excited about. Feng Zhang, who co-invented CRISPR-Cas9 gene editing, called it “genuinely intriguing.” Other scientists, per Bloomberg’s coverage, pushed back that Anthropic oversold the significance.
Both things can be true at once. The funnel design is a legitimately useful pattern. The confidence in the result is a separate question, and it’s the one I want to spend this post on, because I’ve been burned by exactly this failure mode myself — just with blog URLs instead of enzymes.
The funnel is the actual engineering story
200,000 → 3,500 → 20 is a progressive-filtering pipeline, and it’s the same shape you’d design for any large, noisy search space where a full evaluation on every candidate is too expensive to run at scale. You don’t send 950 agents at full analytical depth against 200,000 candidates — you’d burn the token budget on the 99% of candidates that were never going to matter. Instead:
- Stage one is cheap and permissive — probably something close to a structural or sequence-similarity filter, built to let false positives through rather than risk dropping a real hit. 200,000 down to 3,500 is roughly a 98% cut.
- Stage two gets more expensive per candidate and more selective — deeper agent analysis against each of the 3,500, narrowing to 20. This is where the bulk of the 210M tokens probably got spent, because 3,500 candidates getting real agentic reasoning applied is a very different cost profile than 200,000 getting a cheap filter pass.
- Stage twenty is presumably where a human scientist actually looks at the output.
This is the same shape as any pipeline()-over-parallel() design: cheap broad pass first, narrowing stages that get progressively more expensive and more scrutinized, parallelized within each stage rather than running everything through every stage at full cost. If you’re building any kind of agent fan-out over a large candidate space — security findings, log anomalies, code smells across a monorepo — this funnel shape is the right instinct. Don’t run your most expensive verification step against your widest candidate pool.
Where I stop trusting the headline
Here’s the problem: “950 agents found something interesting” and “950 agents correctly found something true” are different claims, and the gap between them is exactly where I got burned on Digest #67 of a content pipeline I run for this blog. I asked a research subagent to cross-check 12 candidate articles against a list of everything already published, and it told me “verified, no overlap” on all twelve. I published eight. Six were duplicates — same story, different outlet, reworded headline. The subagent didn’t lie. It just didn’t actually run the check it claimed to run, or ran it against stale context, and I had no independent verification step to catch it before the content went out.
Scale that failure mode up to “did this agent fleet correctly identify a genuinely novel biological mechanism, or did it find a plausible-looking pattern match that a domain expert would reject on closer inspection.” The second thing is exactly the shape of error LLM-driven research is most prone to: finding something that pattern-matches “interesting” without the causal or mechanistic grounding that would make it actually true. That’s precisely the gap between “Feng Zhang called it genuinely intriguing” and “other scientists said Anthropic oversold it.” Nobody’s claiming fraud. The disagreement is about how much weight an agent-driven discovery claim deserves before independent wet-lab verification — exactly the same category of disagreement I’d want applied to any AI-fleet finding before I acted on it.
The actual lesson for anyone running fan-out at scale
If you’re using agent fan-out for anything where the output drives a real decision — a security audit, a go/no-go on a migration, a scientific claim — the funnel needs a verification stage that is structurally independent from the stage that generated the finding. Not “ask the same model to double-check its own answer.” A genuinely separate check: a different model, a different prompt angle, or — ideally — a human or a deterministic tool that can’t be talked into agreeing with a plausible-sounding story.
Anthropic’s own framing calls this a “verification program,” which is the right instinct; the question is whether the program runs before the press goes out or after. 950 agents narrowing 200,000 candidates to 20 is a pattern worth copying. Trusting the narrowing as proof of truth, without a verification stage that doesn’t share the same blind spots as the fleet that produced the claim, is the mistake — and it’s the same mistake regardless of whether the output is a CRISPR candidate or eight blog posts you were about to publish.



Discussion
Comments are reviewed before publication. Your email is kept private.