A blocked crawler means you cannot be cited by that engine, no matter how good your content is. The two places this happens are robots.txt and your firewall — and the second one is the one nobody checks.
Which user-agents do AI search engines send?
- OpenAI:
GPTBot(training),OAI-SearchBot(search index),ChatGPT-User(live user-triggered fetch) - Anthropic:
ClaudeBot,Claude-User,Claude-SearchBot,anthropic-ai - Perplexity:
PerplexityBot,Perplexity-User - Google:
Google-Extended(AI grounding;Googlebotis separate) - Others:
Bytespider(ByteDance),CCBot(Common Crawl),meta-externalagent(Meta),Applebot-Extended(Apple)
The *-User agents are worth separating from the rest: they fetch a URL because a person just asked about it, so blocking them changes what an individual user sees right now — not just what gets indexed later.
How do I allow them in robots.txt?
Default behaviour already allows everything, so most sites need no change. The failure mode is the opposite one: an inherited or copied rule that blocks them. Be explicit so the intent is visible and reviewable.
# AI search crawlers - explicitly allowed User-agent: GPTBot Allow: / User-agent: OAI-SearchBot Allow: / User-agent: ClaudeBot Allow: / User-agent: PerplexityBot Allow: / User-agent: Google-Extended Allow: / # Everything else User-agent: * Allow: / Disallow: /private/ Sitemap: https://your-domain.com/sitemap.xml
Keep the file in one place. Pointing at a separate sitemap costs nothing and is frequently forgotten.
Why does my firewall return 403 to AI crawlers?
Bot-management features are the usual culprit, not your robots.txt.Challenge-based bot protection scores a request and may issue a JavaScript challenge to anything that is not a browser. Server-side crawlers cannot solve a JS challenge, so they receive a challenge page or a 403 and record a fetch failure.
Check, in this order: bot-fight mode or equivalent, security level set to an under-attack profile, custom WAF rules matching on user-agent, and rate-limiting rules that treat a crawler burst as an attack.
Should I block AI training crawlers?
It is a legitimate choice, and it is separate from search visibility. Blocking CCBot or GPTBot affects model training and grounding; blocking OAI-SearchBot or PerplexityBot affects whether you can be cited in answers at all. Decide which one you are actually objecting to, and never block Googlebot or Bingbot by accident while doing it.
How do I verify the crawlers are not blocked?
- Fetch your own
robots.txtand confirm no group blocks the agents above. - Request a page with an AI crawler user-agent from a server that is not a browser, and check the status code is
200. - Look for the
cf-mitigatedresponse header, which indicates a challenge was issued. - Review your edge security event log for blocked requests from Google, OpenAI or Perplexity address ranges.
The LLMention scanner runs steps 1 to 3 automatically and reports which agents are blocked and why.