Tuesday morning, our LLM cost dashboard showed the exact same number it showed last week. $2 per million input tokens, $10 per million output, $0.20 per million cache reads. Same line, same color, same everything. If I’d stopped there I would’ve filed Sonnet 5.5 under “nothing to see, move on.”
That would’ve been wrong. Anthropic shipped Sonnet 5.5 on September 28, the second model in the 5.5 family after Opus 5.5 landed a week earlier. Pricing didn’t move at all. And yet Anthropic is claiming up to 30% cheaper per task, running 30%+ faster than Sonnet 5. Those two facts only make sense together if the model is doing the same job in fewer tokens and fewer tool calls. Which is exactly what’s happening. Terminal-Bench 4.0 score went from 10.3% to 70.6%. That’s not a small tuning bump — that’s a different category of agentic reliability.
Here’s the part that should bother anyone running an AI cost model at their company: your dashboard cannot see this. Not because the data isn’t there — because you’re not collecting it.
The metric blind spot
Most FinOps setups for LLM workloads track exactly one thing well: dollars per token, usually rolled up by model and by day. It’s an easy number to get, because the API bill hands it to you. So teams build their whole cost story around it.
The problem is agentic workloads don’t spend money per token. They spend money per task — and a task’s cost is a function of how many tokens it took to get there, which is itself a function of how many steps, retries, and tool calls the model needed. A model that’s 20% more expensive per token but finishes a task in half the tool calls is cheaper in every way that matters. A model that’s the same price per token but suddenly needs 3x the steps to finish is a silent regression that your dashboard will report as “no change” for weeks.
Sonnet 5.5 is the first time I’ve seen this gap actually documented by a vendor instead of just theorized by engineers annoyed at their AWS bill. GitHub ran early tests putting it into GitHub Copilot (GA the same day, September 28) and reported it “matched Claude Sonnet 5 on coding tasks while using significantly fewer steps, tokens, and tool calls.” Zendesk’s Abhinay Kathuria said tickets got processed 20% faster in testing. Neither of those is a token-price story. Both are a tokens-per-task and steps-per-task story, and if that’s not a line on your dashboard, you’re going to read this release as a rounding error.
A worked example
Say your team has a CI agent that triages failing test runs — reads the failure, greps the codebase, checks recent commits, decides if it’s a flaky test or a real regression, files or updates a ticket. On Sonnet 5, that task typically runs:
- 14 tool calls (file reads, greps, git log, ticket API)
- 38,000 tokens total (input + output, including retries)
- Cost at $2/$10 per million: roughly $0.29 per triage run
Run the same task on Sonnet 5.5 with the efficiency gains Anthropic and GitHub are describing — call it a 25% reduction in tool calls and tokens, landing inside their claimed 30% band:
- 10-11 tool calls
- ~28,500 tokens
- Cost: roughly $0.22 per triage run
Same $/token. Nothing changed in your pricing tier. But if you run this agent 2,000 times a month across your test suite, that’s the difference between $580 and $440 — a real $140/month saved, invisible to a dashboard that only tracks the token-price line, because the token price literally didn’t move. Multiply that across every agentic workflow your org runs and the invisible number gets a lot less trivial.
What to actually instrument
You need two numbers next to your token cost, not instead of it: steps per task and tokens per task, both bucketed by task type and by model version. Here’s roughly what that looks like as a wrapper around your agent loop — the kind of thing you’d actually paste into a PR, not a full observability platform:
type AgentRunMetrics = {
taskType: string;
model: string;
stepsPerTask: number;
tokensPerTask: number;
toolCallsPerTask: number;
wallClockMs: number;
succeeded: boolean;
};
async function runInstrumentedAgent(
taskType: string,
model: string,
agentLoop: () => AsyncGenerator<AgentStep>
): Promise<AgentRunMetrics> {
const start = Date.now();
let steps = 0;
let tokens = 0;
let toolCalls = 0;
let succeeded = false;
for await (const step of agentLoop()) {
steps += 1;
tokens += step.inputTokens + step.outputTokens;
if (step.isToolCall) toolCalls += 1;
if (step.isFinalAnswer) succeeded = true;
}
const metrics: AgentRunMetrics = {
taskType,
model,
stepsPerTask: steps,
tokensPerTask: tokens,
toolCallsPerTask: toolCalls,
wallClockMs: Date.now() - start,
succeeded,
};
// Ship to whatever you already use — Datadog, a Postgres table,
// a plain log line your pipeline parses. The point is: these fields
// exist and are queryable per model version, not just $/token.
emitMetric("agent.run", metrics);
return metrics;
}
Then the query you actually want to run once a month is: for each taskType, plot tokensPerTask and stepsPerTask by model over time. When you swap Sonnet 5 for Sonnet 5.5, or when the next model drops, that chart moves even if the price-per-token chart doesn’t. That’s the signal. The token-price chart is just an input to a bill, not a measure of what the model is actually costing you to get work done.
Cost per token is the wrong headline metric now
I’ll say it plainly: teams still running their entire AI FinOps story off $/token are measuring the easy thing, not the right thing. It made sense when models were mostly single-shot completions — one prompt, one response, cost is trivially tokens times price. Agentic workloads broke that assumption a while ago, and most cost dashboards haven’t caught up.
What should replace it is cost per completed task, with steps-per-task and tokens-per-task as the diagnostic layer underneath it — the numbers that tell you why the cost-per-task number moved. Anthropic clearly built Sonnet 5.5 with this framing in mind; the GDPval-AA v2.1 score (1844 Elo, nearly matching Opus 5.5’s 1846) tells you they’re optimizing for task completion quality at a lower operational cost, not squeezing the price-per-token lever at all.
If your team ships an agent to production without a steps-per-task and tokens-per-task chart sitting next to the token-price chart, you’re flying blind on exactly the axis models are now competing on. Sonnet 5.5 is a good excuse to go add it before the next release makes the gap even harder to explain to whoever owns the budget.