What we found
Not one of the 30 homepages scored an A. The highest was salesforce.com at 83, and only 2 sites reached a B. Of the 22 that answered with a normal page, the average was 62 out of 100 — a D on this scale.
The second finding is sharper. 8 of the 30 domains disallow at least one of GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot or Google-Extended at their site root. Every news publisher in the sample is in that group: 6 of 6. The publishing industry has not drifted into blocking AI crawlers by accident — it has chosen to.
The third finding is about the companies selling AI. OpenAI scores 55 and Anthropic 62, both below the average of the sample, and neither declares an Organization entity. The sites that score best are the ones selling infrastructure to developers, not the ones selling the answers.
Why the low scores matter more than they look
These are homepages, not articles, and homepages are the hardest page on any site to score well on: they are short, they are navigational, and they carry little of the evidence a generated answer needs. A low score here does not mean a site is invisible in AI answers — their articles may be in far better shape.
What it does mean is that the page an AI system lands on first gives it almost nothing to quote. No statistics, no quotation, usually no question-shaped heading and often no entity markup that says which organisation the domain belongs to. On 22 of the 30, a model reading the homepage has no structured statement of who owns the site.
Blocking is not the same as being refused
Two different things happened to this scan and they are easy to confuse, so the table separates them. A blocked mark means robots.txt disallows an AI agent at the root — a published, deliberate policy that any crawler will obey. A non-200 status means our crawler was turned away, which usually means bot management decided a small scanner from a datacentre looked like a scraper.
Those are not the same claim, and only the first one is evidence about AI access. A site can refuse us and still welcome GPTBot; 4 domains in this table are exactly that case.
The results
How this was done
- The 30 domains were chosen by hand to cover search, AI labs, enterprise software, infrastructure, developer platforms, news and German industry. This is a convenience sample, not a random one, and it is not a ranking of the largest sites on the internet.
- Each homepage was fetched once, on 2 October 2026, by the LLMention crawler from its own infrastructure, identifying itself honestly. It did not impersonate an AI crawler and it was not sent from an AI crawler’s address.
- robots.txt, llms.txt and sitemap.xml were requested from the same origin, and the 38 documented checks were run over the result.
- No AI engine was queried. Nothing on this page is a measurement of what ChatGPT, Claude or Perplexity actually say — that is a different experiment and this tool does not run it.
Limits of this study
- Homepages only. A site’s articles are usually in better shape than its homepage, so these scores understate most of the sites listed.
- One moment in time. robots.txt and markup change constantly; the table is a snapshot, and every row can be re-run.
- Geo-dependence. Sites that redirect by visitor location returned a regional version, and the regional version differs. stripe.com returned its Netherlands edition during one run and its Singapore edition during another.
- A refused request proves nothing about AI access. The scanner fetches as itself. It cannot see what a site does to a real GPTBot, and it does not claim to.
- Not a prediction. These are readiness checks, not a measurement of citations. A site can score 12 and still be quoted tomorrow.
Reproduce it
Every row above comes from one public endpoint, and brief=1 returns just the headline numbers:
curl "https://geo-scanner.ccie13192.com/api/scan?domain=openai.com&brief=1"
{"domain":"openai.com","status":200,"score":55,"grade":"F",
"checksRun":38,"checksPassed":23,"aiCrawlersBlocked":false,
"hasJsonLdEntity":false,"hasSameAs":false,"hasRobotsTxt":true,
"topIssue":"jsonld-valid"}The rules behind every number are published on the methodology page, including the parts of the picture the score cannot see. If a result here looks wrong, that is a bug report and it is welcome — the check will be fixed rather than the number.