THE CRUNCH

Cloudflare has revealed it used frontier AI models to test its own Web Application Firewall (WAF) against a range of attack vectors. The security firm built a system where an LLM acted as a hacker, iterating on payloads to see if the WAF would block them. The tester had no visibility into the WAF's rules or source code, only seeing HTTP responses. The experiment ran 1,107 attempts across six categories, including SQL injection and cross-site scripting, against an authorised staging environment. The vast majority of attacks were blocked, but the few that got through helped Cloudflare create new detections to harden the system for all customers.

The LLM tester operated in an adaptive loop, suggesting variations of an attack based on the WAF's response. It could change encoding, move the payload to different parts of the request, or switch to a different vulnerability. The system was implemented in Python to handle HTTP replay and orchestrate the scenarios. Cloudflare notes that a payload bypassing the WAF still needs an exploitable application to succeed, so keeping software patched remains a critical defence.

The test configuration used Cloudflare's WAF with a blocking score of 30 or below and enabled managed and OWASP Core Rulesets. The six attack categories tested were cross-site scripting, SQL injection, command injection, server-side request forgery, path traversal, and Log4j. Cloudflare emphasises that the results describe the configured WAF boundary as a whole, rather than individual rule performance.

Cloudflare plans to make this testing process a foundational building block of its WAF development lifecycle. The company also offers guidance on correctly deploying a WAF and patching software to reduce the attack surface.