JOURNAL

Cloudflare Built a Security Agent That Checks Its Own Citations Before an Analyst Sees Them

Cloudflare's new Managed Defense harness admits its first prototype hallucinated, then redesigned around code-verified citations instead of trusting one big agent.

Read with AI

Choose content to copy and paste into your AI assistant. Nothing is sent automatically. CMS content is converted to Markdown; original Markdown is used when available.

On October 7, 2026, Cloudflare engineers Deanna Tran and Blake Darché published a blog post with an admission you don’t see often from a vendor: their first version of an AI security analyst hallucinated. Not occasionally. Enough that they threw out the design and rebuilt it around a different principle — the model doesn’t get to assert anything a human can’t trace back to a specific piece of evidence.

The system is an agentic harness inside Cloudflare’s Managed Defense product, currently in early beta and gated to enterprise customers through their account team. It’s not self-serve, and it’s not the same thing as Clef — the open-source decision-model family Cloudflare shipped the same week. Clef is a general-purpose classifier you can run yourself on Workers AI. The SOC harness is a specific internal application that happens to use Clef as one component among several. Worth separating those two, because Cloudflare’s own Workers AI changelog lists a ticket-routing example that explicitly runs “without a human in the loop” — a completely different use case with a completely different risk profile than a security investigation. The SOC harness does not work that way, and conflating the two would misrepresent both.

Recon first, inference second

The architecture starts somewhere unglamorous: before any LLM call happens, deterministic code runs a fixed set of reconnaissance workflows against versioned APIs. No model decides what to look up first. A script does. Only once that evidence exists does Clef get involved — twice, in fact. First as a noise filter on the raw recon output, then again later to score whether the gathered evidence is actually sufficient before the expensive reasoning steps start.

This ordering is the actual design lesson here, more than any specific tool choice. Most agent pipelines I’ve seen reach for a model immediately because that’s the fun part to build. Cloudflare’s team did the boring, auditable part first and only let inference touch data that a script had already collected through a fixed, reviewable path. If something goes wrong, you can replay the recon step independently of anything the model said.

Why a single agent fails

Here’s the line from the post that’s worth sitting with: “Our first prototype showed the limits of one general-purpose agent. It produced useful analysis, but it also hallucinated claims the evidence did not support.” That’s a team admitting their first instinct — one capable model doing the whole investigation — didn’t hold up under the thing security work actually needs, which is claims that survive an audit, not claims that merely sound plausible.

The fix wasn’t a bigger model or a longer system prompt. It was decomposition plus a verification layer that isn’t itself a model. Four specialist agents run in parallel, each scoped narrowly: traffic analysis, customer context, global telemetry, threat intelligence. None of them writes the final advisory. A separate synthesis agent combines their typed findings — and Cloudflare built it so that agent literally cannot fetch new evidence or pick a classification outside an approved vocabulary. It can only arrange what the specialists already found.

Then comes the part that actually answers the hallucination problem: before any of this reaches a human, application code — not a model — checks that every citation in the output exists, belongs to this specific investigation, and actually supports the claim attached to it. If a citation fails that check, the claim doesn’t ship. That’s the whole trick. You don’t ask a second LLM to grade the first LLM’s homework and hope it’s more honest. You write a deterministic checker for the one thing that actually matters: does this citation point at something real.

Flow diagram showing deterministic recon feeding Clef triage, which feeds four parallel specialist agents (traffic, customer, telemetry, threat intel), which feed a synthesis agent, which passes through a code-based citation check before reaching a Managed Defense analyst who owns the final decision.
Architecture described in Cloudflare’s October 7, 2026 blog post on the Managed Defense agentic security harness, currently in early beta.

Advisory, not autonomous — and that’s stated plainly

Cloudflare is explicit that this is advisory only: “The harness produces an advisory and does not act on behalf of Managed Defense Analysts,” and “the Managed Defense Analyst remains responsible for the decision and any mitigation.” No numbers are published on how much time this saves analysts or how often the citation check actually catches a bad claim before it ships — Cloudflare didn’t disclose either, and I’m not going to guess at them. What they did disclose is the design reasoning, which is more useful than a cherry-picked accuracy stat would have been anyway.

What I’d steal from this for other agent pipelines

If you’re building anything where an agent’s output needs to survive scrutiny — compliance, incident response, financial reporting, legal review — the pattern here generalizes past security. Separate evidence collection from reasoning and make the collection step deterministic and replayable. Don’t let your synthesis step invent new facts; constrain it to arranging what earlier steps already verified. And build the citation check as plain application code, not as another model call, because a model checking a model just moves the hallucination risk one step downstream instead of removing it. The part of this architecture that’s actually interesting isn’t that Cloudflare used four agents instead of one. It’s that they didn’t trust any of the agents to grade their own work.

Discussion

Comments are reviewed before publication. Your email is kept private.

← Back to allĐọc tiếng Việt