Cloudflare published a post on September 29 describing something most security teams talk about wanting to build and rarely actually ship: an LLM that attacks your own WAF, reads the response, and mutates the payload for the next attempt. No access to the WAF’s rules. No internal details. Just a black box and a model deciding what to try next.
The numbers: 1,107 mutation attempts across 45 scenarios, six attack categories — XSS, SQL injection, command injection, SSRF, path traversal, Log4j. 558 requests got blocked. 49 findings needed remediation, and 48 of those 49 clustered in just two categories: command injection and SSRF. Three new detections shipped to the managed ruleset, including one specifically for SSRF payloads hiding behind non-standard numeric IP formats. Every single AI-generated finding went through human review before anything shipped — Cloudflare says that explicitly, and I believe them, because model output alone isn’t a finding, it’s a claim.
Good result. Not the interesting part.
The number everyone will skip past
XSS, path traversal, SQL injection, Log4j — strong coverage, basically clean. Command injection and SSRF — 48 gaps between them. Same testing budget, same model, same iteration loop. The difference wasn’t effort. It was that two categories happened to have more genuinely novel mutation space (encoding tricks, numeric IP representations, shell-metacharacter combinations the WAF’s training data hadn’t seen) and the other four didn’t have much room left to surprise anyone.
Here’s the part that should bother you if you run a security test suite: most of us still report coverage as attempt count, or scenario count, or “we ran 1,107 payloads against the API this week.” That number tells you almost nothing about whether you found the gap that mattered. Cloudflare’s own data says the finding rate wasn’t uniform across categories — it was almost entirely concentrated in two of six. If your testing effort is also spread uniformly across categories, and the real risk isn’t, you’re allocating your best asset — model calls, or tester-hours, whichever you’re spending — exactly wrong.
What an attempt-count-blind harness actually looks like
The naive version of this loop just round-robins categories or picks randomly:
categories = ["xss", "sqli", "cmdi", "ssrf", "lfi", "log4j"]
def naive_next_category(attempt_count):
return categories[attempt_count % len(categories)]
Every category gets an equal share of the budget, forever, regardless of what’s actually turning up findings. That’s how you spend 900 of your 1,107 attempts confirming things you already knew were solid.
A diversity-aware version tracks finding rate per category and re-weights toward whatever’s producing signal, while never fully starving a category that’s gone quiet — because “quiet” and “solved” aren’t the same thing:
import random
def next_category(history):
# history: {category: {"attempts": int, "findings": int}}
weights = {}
for cat, stats in history.items():
rate = stats["findings"] / max(stats["attempts"], 1)
# floor weight keeps every category alive; no category ever hits zero
weights[cat] = max(rate, 0.05)
total = sum(weights.values())
r = random.uniform(0, total)
upto = 0
for cat, w in weights.items():
upto += w
if upto >= r:
return cat
return list(history.keys())[-1]
This is a five-minute bandit algorithm, not novel research — the point isn’t the code, it’s that Cloudflare’s raw numbers are a clean argument for why you’d bother writing it. If two categories out of six are producing 98% of your findings, a testing harness that doesn’t know that is leaving the other 900 attempts on the table doing very little.
The human review step is not a footnote
I want to sit on one detail Cloudflare mentioned almost in passing: every AI-generated finding still went through a human before it counted as a finding. It would’ve been easy to skip that line in the writeup — the automation is the impressive part, the review step sounds like boring process. It’s not boring, it’s the whole reason the 49 number is trustworthy at all. An LLM attacking a black box will produce plenty of requests that look like a bypass and aren’t — a WAF returning a 200 because the payload happened to hit a route that was never protected in the first place isn’t a gap, it’s a false positive with good production values. Skip the review step to save time and your finding count goes up while your actual security posture stays exactly where it was, or gets worse because someone ships a “fix” for a bug that was never real.
Where I’d actually apply this tomorrow
Not just WAF testing. Any place your team runs automated security or quality scanning with a fixed budget — fuzzing, dependency scanning, even LLM-based code review bots hunting for bugs — probably reports results the same flat way Cloudflare’s raw attempt count would have, if they’d stopped at “1,107 attempts, 49 findings” and called it done. Break your own numbers down by category before you trust the aggregate. If one category is silent, that’s either good news or a sign your harness never learned to look there — and from the outside, those two look identical until you go check. And keep a human in the loop on anything the model calls a finding, not because the model is untrustworthy, but because “looks like a bypass” and “is a bypass” are a different claim, and only one of them is worth a ticket.