THE CRUNCH

Cloudflare has launched a new setting called Disallow AI Training that lets website owners publish a robots.txt directive to refuse AI model training without blocking search engine crawlers. The feature is designed to solve the tradeoff where refusing AI training also blocks search indexing, a problem caused by mixed-use crawlers that perform both tasks. The company says Apple, Google, and Microsoft have agreed to a

Accountable

designation that respects the new preference. Cloudflare classifies bots by behaviour, separating Search, Training, and Agent activities. The new setting is available at the domain level and applies to all training crawlers, including those run by Amazon, Anthropic, Meta, and OpenAI, while still allowing Accountable mixed-use crawlers to index the site for search.

Cloudflare is also working on controls for AI summaries, which it describes as the next challenge after training. The company aims to let site owners set how much of their content is included in AI summaries directly with operators or through Cloudflare by early next year. This follows a broader trend of website owners seeking granular control over how their content is used by AI systems, rather than a one-size-fits-all block.

The new setting is part of Cloudflare's broader bot management tools, which previously lacked a way to block training-only crawlers without also blocking search crawlers. The company argues that robots.txt directives alone are insufficient because they cannot identify who is crawling or stop a crawler that ignores them. By publishing the preference and identifying the operator, Cloudflare aims to provide transparency and enforcement for site owners.

WHAT HAPPENS NEXT

Cloudflare plans to introduce controls for AI summaries by early next year, allowing site owners to specify how much of their content is included in AI-generated overviews.