Cognition released SWE-2 on September 10, its newest coding model for Devin, and the headline number is a familiar shape: matches a bigger model on the main benchmark, costs a fraction as much. What’s actually interesting isn’t the leaderboard position — it’s the training method behind the cost curve, because it changes how I’d think about picking an effort level for a given task rather than just picking a model.
The number that matters isn’t the top-line score
SWE-2 is post-trained from Kimi K3, a 2.8 trillion-parameter base model, via reinforcement learning. On FrontierCode 1.1 Main — Cognition’s primary coding benchmark — SWE-2 scores 50.0%, essentially tied with Fable 5.1’s 50.9% and behind GPT-6 Astra’s 53.3%. On its own, that’s an unremarkable middle-of-the-pack result.
The number that actually matters is next to it: SWE-2 gets that score at roughly 64% lower cost than Fable 5.1. And against its own predecessor, SWE-1.7, the medium-effort tier needs 58% fewer turns and costs 81% less, with mean steps per run dropping from 127 to 53. On Terminal-Bench 2.1 it actually leads the table outright at 92.8%, ahead of both Fable 5.1 (91.4%) and GPT-6 Astra (89.9%) — though it falls well behind both on the harder Terminal-Bench 4 (27.3% versus 55.8% and 57.9%), which tells you the efficiency gains aren’t free everywhere; they’re concentrated on tasks the model has been specifically trained to solve efficiently, not on the hardest long-horizon problems.
That asymmetry is the real story. This isn’t “smaller model, slightly worse, cheaper” — it’s “similar quality on the tasks it’s tuned for, cheaper by training the model to stop taking unnecessary turns,” and noticeably weaker on the tasks it isn’t tuned for. Those are different engineering claims and worth treating differently when you’re deciding whether to route work to it.
How they actually did it: one RL run, not three models
The mechanism is worth understanding because it’s a pattern I expect other coding-model vendors to copy. Instead of training separate models for “fast and cheap” versus “slow and thorough,” Cognition trained a single RL run that optimizes multiple selectable reasoning-effort levels — medium, high, max — simultaneously, using a cost-penalized reward function of the form R = S − λC, where success is traded off against cost with an effort-level-specific penalty weight.
Concretely, that means the model itself learns different stopping behavior at each effort tier, rather than a product layer bolting a turn limit onto a model that doesn’t know it’s being cut short. In my experience, the difference between those two approaches shows up exactly where you’d expect: a model with an artificial turn cap tends to produce truncated, half-verified work when it hits the limit, while a model trained to know its budget tends to front-load the cheap, high-signal moves — read the right files, form a plan — before spending turns on execution. Cognition’s own numbers back this up indirectly: median steps to first edit dropped from 48 to 18, which reads as the model getting to a concrete action faster rather than just doing less work overall.
What this means if you’re routing work across models
If you’re running a coding-agent pipeline that routes tasks by complexity — and by September 2026 most serious internal tooling setups I’ve seen do this in some form — SWE-2’s medium tier is a genuinely different cost point for the “well-specified, contained” bucket: bug fixes with a clear repro, small refactors, test-writing against an existing pattern. That’s the bucket where FrontierCode-style tasks live and where the 64%-cheaper number is real.
For the “large, ambiguous, multi-file, needs actual judgment” bucket — the Terminal-Bench 4 territory — the data says don’t reach for SWE-2’s efficiency as your primary selection criterion. Route those to whichever model wins on the harder benchmark for your workload, and treat cost there as secondary. A cost-optimized medium-effort model saving turns on a task it’s structurally weak at doesn’t save you money; it just fails cheaper and faster, and you pay the real cost in review time and rework.
One vendor-lock caveat worth flagging to anyone evaluating this for a team: there are no open weights and no standalone API — SWE-2 runs only inside Devin (Desktop and CLI at launch, Web and Fusion rolling out). If your evaluation criteria include “can we self-host this” or “can we call this from our own orchestration layer,” this doesn’t clear that bar regardless of the benchmark numbers, and that’s a decision to make explicit before your team gets attached to the cost curve.
The practical takeaway
The pattern to watch for, beyond this specific release: “one training run, multiple effort tiers with different cost-performance tradeoffs baked in” is a more capital-efficient way for a vendor to cover a product’s low-end and high-end use cases than shipping and maintaining separate model lines. If you’re evaluating coding agents for a team, the question to ask a vendor now isn’t just “what’s your best benchmark score” — it’s “show me the cost-performance curve across your effort tiers, and which of your benchmarks stress the top of that curve versus the bottom.” SWE-2’s own numbers are the clearest illustration yet of why that second question matters: the model that wins on cost and the model that wins on hard tasks are, right now, not the same tier of the same model.