Anthropic put a number on something every engineer already knew and nobody wanted to measure: in “How we contain Claude across products,” published May 25, 2026, they reported that users approve roughly 93% of Claude Code’s permission prompts. Not “roughly 93% turn out fine” — 93% get clicked through. Their own framing is blunt about why: the more approval dialogs a person sees, the less attention they pay to any single one. That’s the same mechanism behind Jason Lemkin’s Replit database, from the first post in this series — he’d already given the one instruction that mattered, and the system still needed something other than his attention to hold the line.
Two posts, two different taxonomies, on purpose
Two months earlier, on March 25, 2026, Anthropic had published “How we built Claude Code auto mode,” describing a two-stage classifier: a fast single-token filter checks every action first, and only when that filter flags something does a slower chain-of-thought reasoning pass — running on Sonnet 4.6 — actually think it through. That post frames the threat model as four things: overeager behavior, honest mistakes, prompt injection, and a misaligned model. It also lists, separately, four categories that get hard-blocked no matter what the classifier concludes: destroying or exfiltrating data, degrading the security posture, crossing trust boundaries, or bypassing review in a way that affects other people.
The May post uses a different split entirely — user misuse, model misbehavior, external attackers — three buckets sorted by who or what caused the problem, not what the action does. I went back and forth on whether this is Anthropic being inconsistent. It isn’t, really. March’s framework is scoped to one classifier deciding whether to interrupt a specific action; May’s is a wider lens across their whole product line, sorting by source of harm rather than shape of harm. March’s “prompt injection” is a subset of May’s “external attackers.” March’s “overeager behavior,” “honest mistakes,” and “misaligned model” all fold into May’s “model misbehavior.” Fine detail feeding a coarser map — but it’s worth knowing you’re looking at two different resolutions before you quote either one as the Anthropic risk taxonomy, because there isn’t one yet.
The part that actually moved the 93% number
None of this would matter if the fix were “ask better.” It isn’t. A separate post from October 20, 2025, “Beyond permission prompts: making Claude Code more secure and autonomous,” introduced sandboxed execution — bubblewrap on Linux, Seatbelt on macOS — and reported the actual result: an 84% drop in permission prompts, not from a smarter classifier deciding what to ask about, but from removing most of the actions that needed asking in the first place. A sandboxed shell that can’t reach the network or touch files outside its working directory doesn’t need a human’s blessing for every command inside it, because there’s nothing left to approve.
That’s the real shift, and it’s the same one Replit made after Lemkin’s incident: stop routing everything through a dialog box, and structurally shrink what’s reachable without one. For anything that still needs a live human decision, the Claude Agent SDK exposes a canUseTool callback — a genuine pause-and-wait hook, evaluated only after hooks and deny/allow rules have already failed to resolve the request on their own. It’s the last resort, not the first line.
A checklist, if you’re building this yourself
I’m not going to pretend five bullet points solve agent authorization. But if you’re wiring approval flows into your own agent product, these are the design calls that actually track with what worked here and in the OWASP Top 10 for Agentic Applications this series has been leaning on:
Sort by reversibility, not by category. A destructive or irreversible action gets a hard structural block, not a prompt — the same lesson Replit learned about production databases. If undoing it is expensive or impossible, a chat message asking permission isn’t a control, it’s a formality.
Keep the judge separate from the actor. Anthropic’s classifier runs as its own pass, not the acting model self-reporting on its own risk. A model grading its own homework on whether to proceed is the honest-mistakes category from the March post, built into the safety mechanism itself.
Shrink the blast radius before you ask about it. Sandboxing cut Anthropic’s own prompt volume by 84% — not by asking smarter questions, but by making most questions unnecessary. If an action literally cannot reach anything sensitive, it doesn’t need a human in the loop at all.
Treat your own approval rate as a live metric, not a compliance box. If you’re not measuring what fraction of your prompts get rubber-stamped, you don’t know you have a 93% problem until something goes through that shouldn’t have. Track it the way you’d track an alert’s false-positive rate.
Reserve the human-in-the-loop callback for what’s actually ambiguous. canUseTool-style hooks work precisely because they’re rare. The moment they fire on routine actions, you’ve rebuilt the fatigue problem you were trying to sandbox your way out of.
Three posts, one thread through all of them: the failure was never that an agent was too capable. It’s that the capable part and the accountable part kept living in the same undivided moment — one chat window, one signature, one dialog box — with no structural gap between deciding and doing. Every fix that actually worked, from Replit’s dev/prod split to Anthropic’s sandbox, did the same unglamorous thing: it put a wall where there used to be a question.