Cloudflare Tests Its WAF Using Frontier AI Models
Cloudflare engineers built a dynamic testing framework driven by frontier AI models to evaluate how effectively their web application firewall handles mutated payloads.

Evaluating WAF Defenses with Frontier AI Models
Security engineers at Cloudflare recently tested their web application firewall by deploying a dynamic testing system driven by frontier AI models. While static security analysis examines code without execution, dynamic methods probe running systems. In this project, the large language models operated without visibility into the source code or internal WAF rules, viewing only selected HTTP response data to simulate an automated attacker.
Large language models excel at rapidly iterating and mutating attack payloads by testing various encodings or changing where payloads appear in HTTP requests. To harness this capability safely, Cloudflare built a Python-based harness rather than wrapping an existing penetration testing framework. The software handled HTTP replay, state tracking, request orchestration, and result collection while ensuring the models could not directly deploy rules or execute network requests.

Structured Testing Across Six Attack Categories
The main evaluation targeted an authorized customer staging environment protected by Cloudflare's security infrastructure. Testers used an allowlisted User-Agent to prevent automated traffic controls from cutting off the test prematurely. The evaluation comprised 45 scenarios, with 44 scenarios focusing directly on six prominent attack categories: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion, and Log4j vulnerabilities. A final separate scenario examined log injection.
During the evaluation, the target zone utilized specific configurations including the OWASP Core Ruleset with Paranoia Level 3, the Cloudflare Managed Ruleset, and WAF Attack Score blocking thresholds set at 30 or below. The system recorded every attempt to measure how effectively the overall boundary handled incoming mutations.
How the Adaptive Testing Loop Operated
The testing architecture functioned through an iterative loop that executed models twice per cycle: first for a proposal call and second for a review call. The proposal call received the starting request, historical context, and previous results to suggest a new variation. Code then validated the target hostname against an allowlist, enforced attempt limits, and dispatched the request.
Following the request, the review call analyzed headers, response status, and response bodies to determine the next predefined step. Because response text could reappear in subsequent prompts, the tester handled all response data as untrusted input. The automated loop terminated when mutations stopped generating useful variations or when the hardcoded attempt limit was reached.

Analyzing Bypasses and Refining Detections
The testing framework generated 1,107 total attempts across the evaluation run. After removing duplicate, malformed, benign, and out-of-scope observations, human reviewers analyzed the remaining leads. While the vast majority of attacks were blocked, specific variations successfully bypassed the security boundary, offering crucial engineering insights.
For instance, in an evaluated SSRF scenario, the tester altered cloud metadata addresses using octal, integer, and trailing-dot representations across different parts of the request. While most variations faced blocks, a trailing-dot representation resulted in a client redirect rather than a firewall block. These findings enabled engineers to develop new detection rules, improving protection for all platform users.

Broader Defense and Future Lifecycle Integration
Cloudflare noted that this adaptive testing methodology is becoming a foundational building block for its ongoing WAF development lifecycle. However, engineers emphasize that a firewall remains only one layer of defense. Because payloads bypassing a WAF still require an underlying vulnerability to succeed, maintaining updated software stacks remains essential for comprehensive security.
Sources
- Cloudflare BlogWe tested our own WAF with frontier AI models. Here’s what we found