Mario Rodriguez, GitHub’s Chief Product Officer, put it plainly in Anthropic’s own release notes: Claude Opus 5.5 “solved more terminal tasks than Opus 5 in less than half the steps.” That’s the kind of quote you only get from someone who ran the model against real workloads before the announcement, not someone reading a spec sheet. Minutes after Anthropic published that quote on September 22, OpenAI shipped two models of its own — GPT-6 Sol and GPT-6 Luna — and neither one is trying to win the same argument.
Two releases, same morning, different math
Anthropic’s pricing move on Opus 5.5 is a straight cut: input tokens dropped from $5 to $4 per million, output from $25 to $20 per million versus Opus 5, and cache reads fell from $0.50 to $0.20 — a 60% cut on the thing that matters most for agent loops that keep re-reading the same context window. Anthropic’s own framing calls it a 40% cost reduction on typical workloads, and the benchmark table backs up why: Terminal-Bench 4.0 went from 52.3% to 66.4%, OSWorld 2.0 from 74.0% to 81.8%, and on GDPval-AA (Anthropic’s agent-quality Elo score) Opus 5.5 landed at 1846 versus Opus 5’s 1708 — ahead of both Fable 5.1 (1735) and GPT-6 Astra (1542) on the same chart.
OpenAI’s answer wasn’t one model, it was two, and the pricing gap between them tells you exactly who each one is for. Sol: $2 input / $10 output per million tokens, aimed at “recurring coding and agent work” — 33.2% on AutomationBench at 27 cents per completed task, 68.8% on software-engineering benchmarks. Luna: ten cents input, fifty cents output — two orders of magnitude cheaper than Sol — built for “high-volume routine jobs such as summarization and extraction.” OpenAI’s own claim is that Luna, run at a higher effort setting, matches its predecessor’s quality “at roughly a hundredth of the cost.” Both models came in at half what their predecessors charged.
Read the shape of the price list, not just the top number
Put the four numbers side by side and the strategic split is obvious: Anthropic cut its flagship’s price by 20-40% and spent the saved margin on fewer steps per task. OpenAI split its lineup into a mid-tier agent model and a near-free bulk-processing model, betting that most production token volume looks like Luna’s job — summarize this, extract that — not like a 38-prompt debugging session. Anthropic is optimizing the expensive 10% of your traffic. OpenAI just made the cheap 90% almost free.
Check this against your own traffic, not the vendor’s benchmark
The benchmark that should actually move your routing config isn’t accuracy — it’s cost per completed task, and that only shows up once you replay your own traced runs through each price point:
def cost_per_task(runs, price_in, price_out):
total_in = sum(r["input_tokens"] for r in runs)
total_out = sum(r["output_tokens"] for r in runs)
return ((total_in / 1_000_000 * price_in) +
(total_out / 1_000_000 * price_out)) / len(runs)
opus_5 = cost_per_task(runs, price_in=5, price_out=25)
opus_5_5 = cost_per_task(runs, price_in=4, price_out=20)
gpt6_sol = cost_per_task(runs, price_in=2, price_out=10)
gpt6_luna = cost_per_task(runs, price_in=0.10, price_out=0.50)
Run your own agent traces through all four price points before picking one model to standardize on. A workload heavy on short extraction calls will make Luna look absurdly cheap and Opus 5.5 look wasteful; a workload that’s mostly multi-step terminal or coding agent work will show the opposite, because Opus 5.5’s fewer-steps-per-task advantage compounds exactly where Luna would need three times the retries to get an equivalent answer.
Where I’d actually route traffic
If your production traffic is a mix — and almost everyone’s is — the honest move isn’t picking a winner, it’s building the router. Send the short, low-ambiguity calls (classification, extraction, summarization) to something Luna-priced, and reserve Opus-tier spend for the calls where GitHub’s own numbers hold: fewer steps, fewer retries, less rework. Clio’s Sean Heintz described the same pattern from the user side — “compared with Opus 5, it hit milestones faster and required minimal reworking” — and Quantium’s Harley Barnes put a number on it: a complex task that used to take 38 prompts over four days came down to 11 prompts over three hours. That’s not a benchmark chart, that’s a team’s actual week getting shorter. The two-tier router that decides which class of call gets which price point is worth more engineering time this quarter than the argument about which lab’s flagship to standardize on.
Sources: Anthropic — Claude Opus 5.5, SiliconANGLE — Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models, The New Stack — OpenAI GPT-6 Sol Luna release