AI Crawler Access

No AI crawler is blocked

One of 40 checks the scanner runs. This one is worth 6 points inside the AI Crawler Access dimension, which carries 16% of the total score.

What does this check look at?

Parsing robots.txt into user-agent groups, none of gptbot, claudebot, perplexitybot, oai-searchbot or google-extended is disallowed from /, AND the homepage served a crawler-shaped request with HTTP 200. An exact agent group takes precedence over the * group, the longest matching pattern wins and Allow wins ties, an empty Disallow value matches nothing, and * is treated as a wildcard. When the homepage refused the request, the verdict follows a second probe of the same URL made with a browser-shaped User-Agent: if that one returns 200 the refusal is aimed at identified crawlers, and the check fails under robots-ai-blocked; if it is refused as well, the refusal is about the scanning address rather than the site, and this check scores 2 of 6 as unverified.

What does a failure mean?

At least one AI crawler is disallowed from the site root, or the server refuses identified crawlers outright.

A site with no robots.txt at all passes this check only when the homepage was actually served, because with no file no crawler is disallowed by it. That is a deliberate reading: absence of robots.txt is a policy gap, not a block, and the gap is reported separately by robots-present. A refusal by the server is the stronger evidence and now overrides the file: robots.txt saying nothing is not permission if the request is rejected before it is read.

How do I fix it?

  • Partial pass: Check in your CDN or WAF whether identified AI crawlers are admitted, and from which networks. This scan cannot settle that from here.

Why is it worth 6 points?

Points are relative weights inside a dimension, and a dimension's weight decides how much of the total it can move. The reasoning for this dimension's weight is published rather than asserted:

DimensionAI Crawler Access — 16% of the total score
Points for this check6 of 23 in this dimension
Worth at most4.2 of the 100 points

Highest weight of the twelve. A page that is blocked or unreadable has no path to being cited, so it dominates the total.

What else is measured in this dimension?

How do I see my own result?

Run a scan on your own domain. The report lists every failed check with the evidence that produced it, so you can tell a real failure from a check that does not apply.

Run a free GEO audit on your own site →

The full method, including the dimensions that are deliberately weighted low and the parts of the picture a single-URL scan cannot see, is on the methodology page.