08 Oct 2026 · 4 min read
ai-agents
Claude Haiku 5.5 and the Real Math Behind Cheap Model Routing
Haiku 5.5 is 75% cheaper and far sharper at agent tasks, but a token-efficiency tax means the real savings aren't the headline number.
Read moreIDEAS, DECISIONS & LESSONS
Notes on AI, architecture and building useful products.
26 articles
08 Oct 2026 · 4 min read
ai-agents
Haiku 5.5 is 75% cheaper and far sharper at agent tasks, but a token-efficiency tax means the real savings aren't the headline number.
Read more22 Sep 2026 · 4 min read
ai
Anthropic and OpenAI released within minutes of each other on September 22 — one chasing fewer tokens per task, the other chasing a hundredth of the cost. Both numbers are real, and they point at different production workloads.
Read more19 Sep 2026 · 5 min read
agents
Anthropic cut Claude Fable 5.1 cache reads 75%. Artificial Analysis found max-effort runs cost 20% more anyway, because output tokens jumped 1.7x. I ran the math on our own agent loop to find out which story was true for us.
Read more15 Sep 2026 · 4 min read
agents
Fable 5.1 can read every earlier Claude model's thinking blocks. No earlier model can read Fable 5.1's. That asymmetry looks like a footnote until your team needs to roll back a production agent mid-incident.
Read more13 Sep 2026 · 7 min read
ai
It's not about which model scores higher. It's about which one still gets the job done when you're not sitting next to it. Notes from a year of switching, paying, and betting real client work on Claude.
Read more05 Sep 2026 · 5 min read
agentic coding
Anthropic shipped a beta that lets you add or remove tools mid-conversation without invalidating the prompt cache. Why that matters for long-running agents, and what I changed in my own agent's tool-loading logic after reading the spec.
Read more02 Sep 2026 · 5 min read
ai
Anthropic didn't touch Fable 5.1's per-token price — it cut cache-read cost by 75%. That's not a discount, it's a signal about which architecture pattern the platform now rewards. Here's how I re-cut my prompt structure around it.
Read more01 Sep 2026 · 5 min read
ai-agents
Claude Sonnet 5's introductory pricing and GitHub Copilot Studio's promotional credits both expired on the same day. Here's the real math on what changed, why it happened simultaneously, and how to budget for the next one.
Read more26 Aug 2026 · 4 min read
ai-agents
Ollama v0.33.0 shipped a UI-driven gateway that lets Claude Desktop run Qwen, DeepSeek, and Kimi models locally — no MCP server, no Anthropic feature, entirely Ollama's own proxy. Here's how it works, the tradeoffs in tokens/sec and tool-calling reliability, and when it's worth flipping on for a team.
Read more