Every GEO check the scanner runs
The scanner runs 40 explicit checks across 12 weighted dimensions. Every one of them is listed here with the exact rule it applies, so a score can be argued with rather than taken on trust. Nothing on these pages is a summary of the method — it is the method.
A check is worth points inside its dimension, and each dimension carries a share of the total. A dimension is not a checklist of equal items: the weights are published below and the reasoning for each one is on its page.
| Dimension | Weight | Checks | Points |
|---|---|---|---|
| AI Crawler Access | 16% | 7 | 17 |
| Machine Readability | 12% | 3 | 12 |
| Content Depth | 11% | 3 | 11 |
| Citability & Evidence | 11% | 4 | 11 |
| Answer Readiness | 10% | 3 | 10 |
| Trust & Authority | 10% | 4 | 10 |
| Semantic Structure | 8% | 3 | 8 |
| Metadata & Discoverability | 7% | 4 | 7 |
| AI Context Files | 5% | 3 | 6 |
| Freshness | 5% | 2 | 5 |
| International Readiness | 3% | 2 | 3 |
| Delivery & Mobile | 2% | 2 | 2 |
AI Crawler Access
16% of totalIf a crawler cannot read the page, nothing else on this list can matter. This is the only dimension where a failure makes the rest of the score moot.
A request for /robots.txt returns HTTP 200 with a non-empty body.
Parsing robots.txt into user-agent groups, none of gptbot, claudebot, perplexitybot, oai-searchbot or google-extended is disallowed from /, AND the homepage served a crawler-shaped request with HTTP 200.
robots.txt contains a Sitemap: line.
The homepage returns HTTP 200 to a non-browser request that identifies itself honestly.
The homepage has no meta robots directive containing noindex.
/sitemap.xml returns a body containing <urlset> or <sitemapindex>.
robots.txt contains a Content-Signal directive with a value on the same line, for example `Content-Signal: search=yes, ai-input=yes, ai-train=no`.
Machine Readability
12% of totalJSON-LD is how a model binds your brand name, domain and product into one entity instead of inferring three unrelated strings.
At least one <script type="application/ld+json"> block parses as JSON and declares at least one @-keyword (@context, @type, @graph or @id).
Parsed JSON-LD contains an @type of Organization, WebSite, Person or LocalBusiness.
Parsed JSON-LD contains an @type of FAQPage, Article, BlogPosting, HowTo, Product, SoftwareApplication or BreadcrumbList.
Content Depth
11% of totalGenerative engines select sources that answer a question thoroughly. Thin pages are rarely retrievable regardless of how well they are marked up.
Visible text, after removing script, style and comment content, contains at least 800 word tokens.
The page contains at least 4 H2 elements.
The page contains both a <ul>/<ol> and a <table>.
Citability & Evidence
11% of totalStatistics, quotations and cited sources are the interventions with the largest measured effect in the published GEO research.
Body text contains at least 5 numeric claims: percentages, currency amounts, multipliers of the form 3.2x, thousands-separated figures, or raw numbers of five digits or more.
The page contains at least one <blockquote> element.
The page links to at least 2 external hosts on the authority list: arxiv.org, doi.org, nature.com, science.org, acm.org, ieee.org, springer.com, sciencedirect.com, any .gov or .edu host, wikipedia.org, github.com, developer.mozilla.org, ...
The page declares a rel=canonical link.
Answer Readiness
10% of totalAnswers are extracted as spans, not pages. A question-shaped heading followed by an immediate answer is the easiest thing for a retrieval system to lift intact.
At least 3 H2 or H3 headings either end with a question mark or begin with how, what, why, which, when, who, where, is, are, do, does, can, should or will.
Parsed JSON-LD contains a FAQPage node.
Of the first 30 heading-then-paragraph pairs (H2/H3 optionally followed by wrapper elements then a <p>), at least 60% have an opening paragraph of 80 words or fewer.
Semantic Structure
8% of totalHeading hierarchy and landmark elements are how a parser locates section boundaries at all.
- 3 ptExactly one H1
The page contains exactly one non-empty <h1>.
At least 3 of main, article, section, header, nav and footer appear as elements.
The raw server response contains at least 100 words of visible text without executing JavaScript.
Metadata & Discoverability
7% of totalTitle, description and canonical control what a search or answer surface can show about the page.
<title> is between 15 and 65 characters long.
The meta description is between 50 and 160 characters long.
At least one meta property beginning with og: is present.
The <html> element carries a lang attribute.
AI Context Files
5% of totalA cheap and optional signal. Google has stated it does not use llms.txt in Search, and crawler support is inconsistent.
/llms.txt returns more than 20 bytes and contains an H1, a > summary line and at least 3 markdown links.
robots.txt is readable, so a crawler policy is published.
The document head contains a <link> with rel="alternate" and type="text/markdown", in either attribute order.
Freshness
5% of totalDated content is deprioritised in generated answers, and an undated page gives an engine nothing to reason about.
A machine-readable date is found in dateModified, datePublished or article:modified_time.
The current calendar year appears in the visible body text.
International Readiness
3% of totalLanguage and region markup decides which language market a page can be retrieved in at all.
At least 2 distinct hreflang language codes are declared.
A language tag carrying a region subtag appears in <html lang>, in og:locale or in an hreflang attribute - for example en-GB, en_GB or de-AT.
Delivery & Mobile
2% of totalA slow or non-mobile-readable page is dropped before any content analysis happens.
A meta viewport tag is present.
The uncompressed HTML response is under 500 KB.
Questions about the checks
Can I see the rule without running a scan?
Yes, and that is the point of this section. Every rule, its point value and the exact condition that makes it pass are published, so you can read the method before deciding whether the score is worth anything.
Are all checks applied to every page?
No. A small number are marked not applicable to particular kinds of page, such as a policy or contact page, and are removed from the calculation rather than counted as failures. Each exclusion is listed on the page for the check it applies to.
What happens when a check cannot be evaluated?
It says so rather than guessing. A refusal aimed at our scanner is reported as unverified, not as a failure of your site, because those are different findings and only one of them is yours to fix.
Run a free GEO audit on your own site →
The dimension weights, the A–F bands and the parts of the picture a single-URL scan cannot see are all on the methodology page.