Last Tuesday a client asked me why their Claude Code session burned 40,000 tokens on a task that should’ve taken 4,000. I didn’t have a good answer. I had logs, sure, but nothing that told me which tool call did it, in real time, before the bill came due.
So when Anthropic quietly shipped Mods this week — a TypeScript hook API that lets you register middleware into the Claude Code event lifecycle — I dropped what I was doing and tried it on that exact client’s repo.
What a Mod actually is
Forget the marketing name for a second. A Mod is a function you register against a named event, and Claude Code calls it synchronously at that point in the loop. The events that matter for most people: session.start, prompt.submit, tool.call, turn.start, turn.complete, command.run, and ui.render. That’s it — seven hooks, not seventy.
The API shape is register(on, options), and if you’ve ever written Express or Koa middleware, your hands already know what to type:
import { register } from "@anthropic-ai/claude-code-mods";
register("tool.call", {
name: "cost-tracker",
handler: async (ctx, next) => {
const start = ctx.timestamp;
const result = await next();
const tokens = result.usage?.totalTokens ?? 0;
if (tokens > 2000) {
console.warn(
`[cost-tracker] ${ctx.toolName} burned ${tokens} tokens in a single call`
);
}
return result;
},
});
Twenty minutes, and I had a mod logging every tool call over 2,000 tokens straight to stderr. The first thing it caught: a Read call on a 40MB log file the agent had grabbed without a line limit, three different times in the same session, because it kept forgetting it had already read it.
That’s not a prompt problem. That’s an observability problem, and Claude Code finally has a seam for solving it without screen-scraping the CLI output like I’d been doing since March.
Where it broke
Here’s the part the announcement post won’t tell you. tool.call fires before the model decides whether to retry a failed tool. I registered a second mod on turn.complete assuming it would give me one clean summary per conversational turn — instead it fired mid-stream whenever the model paused to “think” between tool calls in a long agentic chain, which meant my cost tracker triple-counted a single logical turn that had four internal tool round-trips.
I didn’t find this in the docs. I found it because my dashboard showed $340 of estimated spend on a session that cost $38 in the actual Anthropic console. Classic middleware bug, the kind anyone who’s debugged a double-fired Express next() has seen before — except here the “request” is a non-deterministic model turn, so you can’t just add an idempotency key and call it done. I ended up keying my dedup logic off ctx.turnId plus a monotonic sequence number Anthropic exposes on the context object, which isn’t documented anywhere I could find — I only noticed it in a console.log(ctx) dump.
The part I actually like
Mods don’t require forking Claude Code or running a proxy in front of the Anthropic API, which is how I was doing cost attribution before — a dumb reverse proxy sniffing request/response bodies, fragile and slow. Now it’s in-process, typed, and ships with the CLI. For a Tech Lead rolling this out across a team, that’s the difference between “here’s a shell script, good luck” and “npm install this, here’s your config.”
I’d still call this a v0.1 surface. The event list is thin — no error, no context.compact hook even though compaction is probably the single most expensive thing that happens in a long session. If you’re building anything that needs to react to context window pressure, you’re stuck polling ctx.contextUsage from inside turn.complete, which is a hack, not a feature.
What I’d ship this week
If you’re running Claude Code for a team bigger than three people, write the cost-tracker mod above today. It’s twenty minutes and it will catch something embarrassing in your first hour — it caught something in mine. Don’t bother with a fancier version until you’ve seen what your actual agents are doing wrong; mine turned out to be a dumb re-read bug, not some elegant architecture problem.
I’m not rewriting my proxy-based setup yet. The turn-counting bug cost me an afternoon, and I want to see at least one more minor version before I trust turn.complete on anything billing-adjacent. But the seven hooks are the right seven hooks, and that’s rarer than it sounds for a v1 API.