# No AI crawler is blocked

How the No AI crawler is blocked check works, what a failure means, and how to fix it in the AI Crawler Access dimension.

- Dimension: AI Crawler Access (16% of the total score)
- This check: 6 points
- Check id: `robots-ai-allowed`

## What does this check look at?

Parsing robots.txt into user-agent groups, none of gptbot, claudebot, perplexitybot, oai-searchbot or google-extended is disallowed from /, AND the homepage served a crawler-shaped request with HTTP 200. An exact agent group takes precedence over the * group, the longest matching pattern wins and Allow wins ties, an empty Disallow value matches nothing, and * is treated as a wildcard. When the homepage refused the request, the verdict follows a second probe of the same URL made with a browser-shaped User-Agent: if that one returns 200 the refusal is aimed at identified crawlers, and the check fails under robots-ai-blocked; if it is refused as well, the refusal is about the scanning address rather than the site, and this check scores 2 of 6 as unverified.

## What does a failure mean?

At least one AI crawler is disallowed from the site root, or the server refuses identified crawlers outright.

> A site with no robots.txt at all passes this check only when the homepage was actually served, because with no file no crawler is disallowed by it. That is a deliberate reading: absence of robots.txt is a policy gap, not a block, and the gap is reported separately by robots-present. A refusal by the server is the stronger evidence and now overrides the file: robots.txt saying nothing is not permission if the request is rejected before it is read.

## How do I fix it?

- **partial** — Check in your CDN or WAF whether identified AI crawlers are admitted, and from which networks. This scan cannot settle that from here.

## Why is it worth 6 points?

Highest weight of the twelve. A page that is blocked or unreadable has no path to being cited, so it dominates the total.

## What else is measured in this dimension?

- [robots.txt is served](https://geo-scanner.ccie13192.com/checks/robots-present/) — 2 points
- [robots.txt declares a sitemap](https://geo-scanner.ccie13192.com/checks/robots-sitemap/) — 1 point
- [Homepage returns HTTP 200](https://geo-scanner.ccie13192.com/checks/home-200/) — 2 points
- [Homepage is marked noindex](https://geo-scanner.ccie13192.com/checks/home-noindex/) — 2 points
- [sitemap.xml is present and well-formed](https://geo-scanner.ccie13192.com/checks/sitemap-valid/) — 3 points
- [robots.txt states a content policy](https://geo-scanner.ccie13192.com/checks/content-signal/) — 1 point

## How do I see my own result?

Run a free scan on your own domain at https://geo-scanner.ccie13192.com/ - the report lists every failed check with the evidence that produced it.

The full method is published at https://geo-scanner.ccie13192.com/methodology/
