A few months ago I spent a week chasing p99 latency on a client’s agent pipeline. Every individual API call looked fine on its own graph — 80ms, 120ms, nothing alarming. Then I added up a full tool-call loop: retrieve, reason, call a tool, reason again, call another tool, reason once more. Six hops, each with its own round trip, and the total was well past a second before the model even started generating the final answer. Nobody budgets for that until they measure it. That’s the frame I read the new Anthropic-Akamai deal through, and it changes what the announcement actually means.
What was actually signed
On September 24, 2026, Akamai announced an $11.6 billion multi-year agreement with Anthropic, running seven years, with an option to expand by up to another $9 billion — so a real ceiling near $20.6B if both sides exercise it fully. The equity side is the part worth reading twice: Akamai issued Anthropic a warrant for 7.7 million shares, roughly 5% of Akamai’s outstanding stock, struck at $111.33 a share. About 2% vests immediately just for signing; the remaining 3% vests in roughly 1% increments for every additional $3 billion of cloud services Anthropic actually buys, up to that $9B expansion cap. That’s not typical customer-vendor paper — it’s a stake sized to reward Anthropic for actually using the capacity, not just committing to it on a press release.
What the release does not say is which workloads go where. No inference-vs-training split, no batch-vs-latency-sensitive breakdown. The language is deliberately generic: supporting “Anthropic’s accelerating CPU workload demand” and helping it “build, deploy, and operate AI workloads at scale.” If you want the technical argument for why this deal makes sense, you have to go read what Akamai’s own CTO wrote five months earlier.
The 28ms argument
Back in April 2026 — right when Anthropic shipped its “Managed Agents” architecture, splitting a stateless reasoning “brain” from an execution “hands” layer that can be spun up and torn down independently — Akamai CTO Robert Blumofe published a post making the distributed-infrastructure case explicit. His number: a user in London hitting a Virginia-based inference endpoint eats roughly 28ms of propagation delay each way before a single token comes back. Not compute time. Just the physics of light in fiber.
28ms one-way doesn’t sound like a lot until you put it inside a tool-call loop like the one I was debugging. If an agent’s reasoning step lives in Virginia and every tool call it makes has to round-trip back through Virginia too — vector search, a database lookup, a call to another service — you’re not paying 28ms once, you’re paying it on every hop, compounding with actual compute time at each stage. Move the execution layer physically closer to where the tool calls happen, and you cut that repeated tax, even if the heavy reasoning still happens on GPUs somewhere centralized.
That’s a genuinely good argument for edge-distributed execution. It is a much weaker argument for edge-distributed inference, and the deal announcement blurs the two on purpose, because “CPU workload” is vague enough to cover request routing, embeddings, tool orchestration, and pre/post-processing — all of which run fine on Akamai’s edge CPU fleet — without committing to move the actual model weights anywhere near an edge node. Frontier model inference is still overwhelmingly GPU-bound and still concentrated in a handful of regions; nothing about a CPU-capacity deal changes that.
My actual read on it
I don’t think this deal is primarily a latency fix, despite the marketing gloss borrowed from Blumofe’s post. I think it’s capacity hedging with a latency story attached, because capacity hedging alone doesn’t make for a good press release and “we bought more CPU because we’re growing fast” undersells a $20B number. The tell is the warrant structure: it’s built to reward actual usage ramping up over years, which is exactly what you’d design if your real concern is “do we have enough compute lined up as demand scales,” not “do we have a latency problem to solve this quarter.”
That doesn’t make the CPU-workload point irrelevant — it’s the opposite, actually. As agent systems mature, the CPU-side work (routing, retrieval, orchestration, tool execution, guardrail checks) grows faster than most teams expect relative to the GPU-side reasoning work, and it’s exactly the kind of workload that benefits from being physically close to users and to the services it’s calling. If you’re architecting an agent platform right now, the useful takeaway isn’t “go multi-region for your LLM calls” — it’s “measure how much of your total latency budget is actually GPU inference versus everything else in the loop,” because for a lot of production agent systems, “everything else” is bigger than people assume, and it’s the part you can actually move.
I went back and re-measured that six-hop pipeline from a few months ago with that split in mind. GPU inference was 40% of total latency. The other 60% was retrieval, tool calls, and network hops — all things a CPU edge fleet, placed well, could shrink. Akamai’s bet might be a better one than the press release lets on.