Cloudflare change shifted which AI bots reach your site
A study of 1,046 sites found search crawlers got in far more often after the change, and training crawlers got blocked slightly more.
Cloudflare, a service many sites use for speed and security, changed a default setting on 15 September, and the effect on AI bots was not what was advertised.
@SeenSureHQ measured crawler access to 1,046 websites before and after the change, across eight crawler identities, with a control group of sites not on Cloudflare. The results were shared by Aleyda Solis.
What moved
On Cloudflare sites, refusals went up a little for the two training crawlers, the bots that collect text to train AI models. ClaudeBot went from 19.9% refused to 22.8%, and GPTBot from 18.9% to 22.0%.
Refusals fell hard for the other six. All three search crawlers and all three agent crawlers dropped by 12.7 to 13.9 points each. The biggest mover was OAI-SearchBot, OpenAI's search crawler, going from 16.9% refused to 3.0%. The smallest of the six was Perplexity-User, 14.9% to 2.2%.
The control group of non-Cloudflare sites stayed flat, which suggests the shift came from the Cloudflare change and not from something else.
According to the post, Cloudflare's own stated scope for the change was training and agent crawlers, with search crawlers expected to hold flat. Search crawlers moved the most of any category, and agent crawlers moved in the opposite direction from what a tighter default would predict.
Why a small business owner should care
If you want AI tools to be able to read and cite your pages, your firewall settings decide that, and you may not have picked those settings yourself. In a reply on the same thread, Cittago said their own site had 79% of AI crawler requests refused by the firewall while Googlebot got through 99% of the time, and that nobody chose that setup on purpose.
If your site sits behind Cloudflare, log in, open the bot and firewall settings, and check which AI crawlers are set to blocked. Decide on purpose which ones you want in. Blocking training crawlers while letting search crawlers through is a normal choice, but make it yourself instead of inheriting a default.
If you do not use Cloudflare, nothing here changed for you.
Our own site had 79% of AI crawler requests refused by the firewall while Googlebot got through 99% of the time. Nobody chose that config on purpose.
