Research

We scanned the homepages of 30 major websites for AI visibility

Every result below came from the same scanner this site ships, run against the homepages of 30 large websites on 2 October 2026. Nothing was adjusted by hand, and every figure can be reproduced from the public API.

62
average score, out of 100
0
sites graded A
8/30
disallow AI crawlers in robots.txt
8
did not return a normal 200

What we found

Not one of the 30 homepages scored an A. The highest was salesforce.com at 83, and only 2 sites reached a B. Of the 22 that answered with a normal page, the average was 62 out of 100 — a D on this scale.

The second finding is sharper. 8 of the 30 domains disallow at least one of GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot or Google-Extended at their site root. Every news publisher in the sample is in that group: 6 of 6. The publishing industry has not drifted into blocking AI crawlers by accident — it has chosen to.

The third finding is about the companies selling AI. OpenAI scores 55 and Anthropic 62, both below the average of the sample, and neither declares an Organization entity. The sites that score best are the ones selling infrastructure to developers, not the ones selling the answers.

Why the low scores matter more than they look

These are homepages, not articles, and homepages are the hardest page on any site to score well on: they are short, they are navigational, and they carry little of the evidence a generated answer needs. A low score here does not mean a site is invisible in AI answers — their articles may be in far better shape.

What it does mean is that the page an AI system lands on first gives it almost nothing to quote. No statistics, no quotation, usually no question-shaped heading and often no entity markup that says which organisation the domain belongs to. On 22 of the 30, a model reading the homepage has no structured statement of who owns the site.

Blocking is not the same as being refused

Two different things happened to this scan and they are easy to confuse, so the table separates them. A blocked mark means robots.txt disallows an AI agent at the root — a published, deliberate policy that any crawler will obey. A non-200 status means our crawler was turned away, which usually means bot management decided a small scanner from a datacentre looked like a scraper.

Those are not the same claim, and only the first one is evidence about AI access. A site can refuse us and still welcome GPTBot; 4 domains in this table are exactly that case.

The results

DomainScoreHTTPAI crawlersEntitysameAs
salesforce.com83 B200permittedyesno
siemens.com80 B200permittedyesno
vercel.com79 C200permittedyesyes
cloudflare.com77 C200permittedyesyes
stripe.com75 C200permittedyesyes
shopify.com72 C200permittednoyes
apple.com69 D200permittedyesyes
developer.mozilla.org66 D200permittednono
bbc.com63 D200disallowednoyes
anthropic.com62 D200permittednono
notion.so62 D200permittednono
figma.com62 D200disallowedyesyes
nvidia.com61 D200permittednoyes
github.com60 D200permittednono
spiegel.de56 F200disallowedyesyes
openai.com55 F200permittednono
telekom.com54 F200permittednono
bahn.de52 F200permittednono
theguardian.com49 F200disallowednono
microsoft.com48 F200permittednono
wikipedia.org47 F200permittednono
sap.com41 F200permittednono
google.com30 F429permittednono
zeit.de26 F403disallowednono
perplexity.ai23 F403permittednono
stackoverflow.com16 F403permittednono
amazon.com14 F202disallowednono
nytimes.com12 F403disallowednono
reuters.com12 F401disallowednono
bmw.com12 F520permittednono

Click a domain to run the scan yourself and see the current result. Scores in this table are a snapshot taken on 2 October 2026 and will have moved since.

How this was done

  1. The 30 domains were chosen by hand to cover search, AI labs, enterprise software, infrastructure, developer platforms, news and German industry. This is a convenience sample, not a random one, and it is not a ranking of the largest sites on the internet.
  2. Each homepage was fetched once, on 2 October 2026, by the LLMention crawler from its own infrastructure, identifying itself honestly. It did not impersonate an AI crawler and it was not sent from an AI crawler’s address.
  3. robots.txt, llms.txt and sitemap.xml were requested from the same origin, and the 38 documented checks were run over the result.
  4. No AI engine was queried. Nothing on this page is a measurement of what ChatGPT, Claude or Perplexity actually say — that is a different experiment and this tool does not run it.

Limits of this study

  • Homepages only. A site’s articles are usually in better shape than its homepage, so these scores understate most of the sites listed.
  • One moment in time. robots.txt and markup change constantly; the table is a snapshot, and every row can be re-run.
  • Geo-dependence. Sites that redirect by visitor location returned a regional version, and the regional version differs. stripe.com returned its Netherlands edition during one run and its Singapore edition during another.
  • A refused request proves nothing about AI access. The scanner fetches as itself. It cannot see what a site does to a real GPTBot, and it does not claim to.
  • Not a prediction. These are readiness checks, not a measurement of citations. A site can score 12 and still be quoted tomorrow.

Reproduce it

Every row above comes from one public endpoint, and brief=1 returns just the headline numbers:

curl "https://geo-scanner.ccie13192.com/api/scan?domain=openai.com&brief=1"

{"domain":"openai.com","status":200,"score":55,"grade":"F",
 "checksRun":38,"checksPassed":23,"aiCrawlersBlocked":false,
 "hasJsonLdEntity":false,"hasSameAs":false,"hasRobotsTxt":true,
 "topIssue":"jsonld-valid"}

The rules behind every number are published on the methodology page, including the parts of the picture the score cannot see. If a result here looks wrong, that is a bug report and it is welcome — the check will be fixed rather than the number.