Technical

Configuring robots.txt and WAF for GPTBot and PerplexityBot

Which user-agents AI search engines send, how to allow them in robots.txt without opening your firewall, and how to confirm that a WAF is not silently returning 403 to AI crawlers.

4 min read · Updated October 2026

A blocked crawler means you cannot be cited by that engine, no matter how good your content is. The two places this happens are robots.txt and your firewall — and the second one is the one nobody checks.

Which user-agents do AI search engines send?

  • OpenAI: GPTBot (training), OAI-SearchBot (search index), ChatGPT-User (live user-triggered fetch)
  • Anthropic: ClaudeBot, Claude-User, Claude-SearchBot, anthropic-ai
  • Perplexity: PerplexityBot, Perplexity-User
  • Google: Google-Extended (AI grounding; Googlebot is separate)
  • Others: Bytespider (ByteDance), CCBot (Common Crawl), meta-externalagent (Meta), Applebot-Extended (Apple)

The *-User agents are worth separating from the rest: they fetch a URL because a person just asked about it, so blocking them changes what an individual user sees right now — not just what gets indexed later.

How do I allow them in robots.txt?

Default behaviour already allows everything, so most sites need no change. The failure mode is the opposite one: an inherited or copied rule that blocks them. Be explicit so the intent is visible and reviewable.

# AI search crawlers - explicitly allowed
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

# Everything else
User-agent: *
Allow: /
Disallow: /private/

Sitemap: https://your-domain.com/sitemap.xml

Keep the file in one place. Pointing at a separate sitemap costs nothing and is frequently forgotten.

Why does my firewall return 403 to AI crawlers?

Bot-management features are the usual culprit, not your robots.txt.Challenge-based bot protection scores a request and may issue a JavaScript challenge to anything that is not a browser. Server-side crawlers cannot solve a JS challenge, so they receive a challenge page or a 403 and record a fetch failure.

Check, in this order: bot-fight mode or equivalent, security level set to an under-attack profile, custom WAF rules matching on user-agent, and rate-limiting rules that treat a crawler burst as an attack.

Should I block AI training crawlers?

It is a legitimate choice, and it is separate from search visibility. Blocking CCBot or GPTBot affects model training and grounding; blocking OAI-SearchBot or PerplexityBot affects whether you can be cited in answers at all. Decide which one you are actually objecting to, and never block Googlebot or Bingbot by accident while doing it.

How do I verify the crawlers are not blocked?

  1. Fetch your own robots.txt and confirm no group blocks the agents above.
  2. Request a page with an AI crawler user-agent from a server that is not a browser, and check the status code is 200.
  3. Look for the cf-mitigated response header, which indicates a challenge was issued.
  4. Review your edge security event log for blocked requests from Google, OpenAI or Perplexity address ranges.

The LLMention scanner runs steps 1 to 3 automatically and reports which agents are blocked and why.