JOURNAL
SIFT: What It Actually Costs to Make a Coding Agent Improve Itself
MIT and Sakana AI's SIFT framework beat the prior self-improving-agent SOTA by using a Bradley-Terry judge instead of more patches. I ran the real cost numbers — $150 and under 5 hours wall-clock for a 4.4-point jump on Polyglot.

On this page
Darwin Gödel Machine was the headline self-improving-coding-agent result for most of this year — an agent that rewrites its own code, benchmarks the rewrite, and keeps the version that scores higher. The obvious next move for anyone copying that idea is: generate more candidate rewrites, benchmark more of them, let the best one win by brute force. MIT and Sakana AI’s SIFT paper does the opposite. It spends almost no extra compute generating patches and instead puts the budget into picking which patch is actually better — and it beats DGM’s score doing it.
That’s the part worth sitting with before the benchmark numbers. The lever that moved the needle wasn’t “more attempts,” it was “a better judge.”
The piece DGM was missing
DGM’s loop is: mutate the agent’s own code, run it against a benchmark subset, keep the mutation if the score goes up. The scoring step is a hard benchmark pass — expensive, slow, and it only tells you the aggregate number, not which specific change helped.
SIFT adds a second signal that’s cheap by comparison: a pairwise LLM judge. Instead of asking “does this patch score higher on the benchmark,” SIFT asks a judge model “between this patch and that patch, which one is actually better,” repeatedly, across pairs. Those pairwise comparisons get fed into a Bradley-Terry model — the same math behind chess Elo and a lot of RLHF reward modeling — which converts a pile of noisy pairwise votes into a single ranked score per candidate. Then SIFT keeps an archive of up to the top 10 performers by that ranking, rather than just the single best-so-far, so a patch that’s strong but not currently #1 doesn’t get thrown away.
Three pieces running async against each other: patch generation, judging, and benchmarking don’t block on each other in lockstep. While one patch is being benchmarked, the judge is already ranking the last batch, and generation is already producing the next one.
The numbers that justified the design
On Polyglot (the harder of the two evals SIFT reports), using o3-mini as the base agent:
DGM (prior SOTA): 30.7%
SIFT, no judge (ablation): 29.8%
SIFT, full pipeline: 35.1%
Read that ablation row carefully — it’s the whole argument. Strip the judge out and SIFT is slightly worse than DGM, not better. The entire 4.4-point gain over DGM comes from the judge, not from generating more patches or running a fancier search. If you only remember one number from this paper, it’s that the judge is 100% of the improvement and the extra generation compute is close to free.
On TerminalBench, the judge’s top pick scored 36.7% versus 28.1% for the best-performing agent when you don’t use the judge’s ranking to select — same story, different benchmark.
And the part I actually wanted to know before writing this: what does a full run cost. For the 50-task Polyglot eval, SIFT reports roughly $150 total spend, about 42 CPU-hours, and wall-clock under 5 hours. Breaking that down further: each individual patch generation runs about 12 cents, and each judge comparison call runs about 4.4 cents. The judge calls are the cheap part per-call — the cost adds up because you’re running a lot of pairwise comparisons, not because any single comparison is expensive.
Why this matters more than the leaderboard number
If you’re building any kind of self-improving or self-evaluating agent loop — not necessarily a coding agent, this generalizes to any setup where an agent produces multiple candidate outputs and you need to pick one — the instinct is almost always to spend more compute on generation. Try more things, sample more, search wider. SIFT’s ablation is a clean counter-example: the bottleneck wasn’t candidate diversity, it was candidate comparison. A single hard benchmark score per candidate throws away almost all the signal about why one candidate beat another, and pairwise comparison recovers some of that signal cheaply.
The Bradley-Terry choice specifically matters because it’s robust to a judge that’s individually noisy on any single comparison. You don’t need your LLM judge to be right every time — you need it to be right more often than not across many pairs, and the ranking math averages out the noise. This is the same reason Elo works for chess ratings built from individually noisy single games.
If I were replicating this for an internal eval loop, the part I’d build first isn’t the mutation/generation step — most teams already have some version of that. It’s the judge-plus-Bradley-Terry ranking layer, because the ablation data says that’s where 100% of the measured gain lives, and it’s reusable across completely different generation strategies.
Where I’d actually use this
Not every team needs a self-rewriting agent — that’s a narrow, research-flavored use case. But the pattern underneath it is not narrow at all: anytime you have an agent producing N candidate solutions (PR descriptions, SQL query plans, config diffs, test suites) and you’re currently picking a winner by a single automated score, a cheap pairwise judge with Bradley-Terry aggregation is probably a bigger lever than generating more candidates. SIFT’s own ablation is the evidence — before you spend more compute making more things, spend a little compute getting better at telling them apart.



Discussion
Comments are reviewed before publication. Your email is kept private.