11 Sep 2026 · 4 min read
agentic coding
Mercury 2.5 and the Diffusion LLM Bet: 1,107 Tokens/Second Changes Agent UX Math
Inception Labs shipped Mercury 2.5 on September 8, hitting 1,107 tok/s on standard NVIDIA GPUs — 10x a typical autoregressive model — by generating text like image diffusion instead of one token at a time. What that means for latency-bound agent architectures.
Read more







