Skip to content
Sandeep Kumar ChaudharySandeep
Back to BlogAI Search

Cloudflare Was Silently Blocking ChatGPT and Claude

By Sandeep Kumar ChaudharyOct 10, 20262 min read
AI assistant reading web pages

TL;DR

Cloudflare's 'Block AI bots' rule returned HTTP 403 to AI crawlers and its 'Managed robots.txt' prepended Disallow rules to my own robots.txt. Neither was visible from the site's code. Turning both off let ChatGPT, Claude, Perplexity and other AI crawlers read the site.

Key takeaways

  • Your robots.txt is not the only thing deciding who can crawl your site. Your CDN can override it.
  • Test crawler access with the crawler's real user agent, through your public domain.
  • Cloudflare's managed robots.txt prepends its own rules above yours.
  • If you want to be cited in AI answers, AI crawlers need to be able to read your pages.

I spent time making this site easy for AI assistants to understand: structured data, an llms.txt file, and a robots.txt that explicitly welcomes GPTBot, ClaudeBot, PerplexityBot and others. Then I tested whether those crawlers could actually reach the site.

They could not. Every one of them got HTTP 403 Forbidden.

How I tested

The simplest test is to request the homepage while pretending to be each crawler, using its published user agent:

curl -s -o /dev/null -w "%{http_code}\n" \
  -A "Mozilla/5.0 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)" \
  https://sandeepkumarchaudhary.com/

I ran this for fourteen crawlers. Googlebot, Bingbot, DuckDuckBot and Applebot got 200. Ten AI crawlers got 403: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Perplexity-User, Amazonbot, Meta's crawler and Common Crawl's CCBot.

My server was not doing this. The same requests sent straight to the origin returned 200. Something at the edge was blocking them.

Culprit 1: "Block AI bots"

In the Cloudflare dashboard, under Security → Settings → Bot traffic, there was a setting called Block AI bots, scoped to "Block on all pages". It deploys a Cloudflare-managed rule that refuses requests from crawlers Cloudflare classifies as AI bots, before they ever reach your server.

Your robots.txt has no say in this. The crawler is turned away at the door.

I changed the scope to "Do not block (allow crawlers)". Cloudflare also asked about "mixed-purpose crawlers", the ones used for both search and AI training, and I chose to keep allowing them.

Cloudflare has since replaced this setting with AI bot policies, which split crawlers into Search, Agent and Training groups. If you see that newer screen, check all three.

Culprit 2: "Managed robots.txt"

Fixing the 403 was not the end of it. When I read the live robots.txt, it did not match the file my site generates. Above my rules was a block marked "BEGIN Cloudflare Managed content":

User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

This comes from AI Crawl Control → Signals → Managed robots.txt. Cloudflare prepends it to whatever robots.txt your site serves. Crawlers read rules from the top, so these Disallow: / lines told AI crawlers to stay out even after the 403 was gone.

I switched Managed robots.txt off. The live file then matched exactly what the site generates, with zero Disallow lines.

The result

After both changes, I reran the same fourteen requests. All returned 200, and the live robots.txt contained only my own rules.

Why this matters

Being cited in answers from ChatGPT, Claude or Perplexity is becoming a real source of visits, sometimes called generative engine optimization, or GEO. It starts with the most basic requirement: the crawler has to be able to read your page. No amount of structured data helps if the request is refused.

Blocking AI crawlers is a legitimate choice, and some publishers make it deliberately. The problem is making it by accident, through a CDN default you never saw.

A checklist

  1. Test each crawler's user agent against your public domain, not just your origin server.
  2. Compare the live /robots.txt with the one your code generates.
  3. In Cloudflare, check both Block AI bots or AI bot policies, and Managed robots.txt.
  4. Retest after every change, and again after any CDN plan or dashboard update.

Your CDN sits between every crawler and your site. Make sure it is following your intentions, not its defaults.

#Cloudflare#AI crawlers#GEO#robots.txt

Frequently Asked Questions

How do I check whether my site blocks AI crawlers?

Request your homepage with each crawler's user agent, for example curl -A with GPTBot or ClaudeBot, and look at the status code. A 403 means something in front of your server is blocking it.

Should I allow AI crawlers?

It depends on your goals. Blocking them keeps your content out of AI training and answers. Allowing them makes it possible for assistants like ChatGPT, Claude and Perplexity to read and cite your pages.

Does allowing AI crawlers affect Google search ranking?

No. Googlebot is separate from AI crawlers. Google-Extended only controls whether Google may use your content for its AI models, not how you rank in search.

Sandeep Kumar Chaudhary

Sandeep Kumar Chaudhary

Full Stack Software Developer· Nepal's SEO, AEO, GEO & AIO expert and share-market educator. More about me