Methodology

How the GEO score is calculated

Every rule on this page is the rule the scanner actually executes. The list below is generated from the same data the analyser runs on, so the published method and the live method cannot drift apart.

How is the total score calculated?

The scanner runs 38 explicit checks across twelve dimensions. Each check is worth a fixed number of points inside its dimension. A dimension’s score is the fraction of its available points that were earned, and the total is those twelve scores combined by weight:

total = Σ ( dimension_score × dimension_weight / 100 )

where dimension_score = earned_points / available_points × 100
and   Σ dimension_weight = 100

There are no score floors. A page that satisfies none of the checks scores near zero rather than being lifted to a respectable-looking minimum. Points are relative weights within a dimension: where two entries below describe the two outcomes of a single check, only one of them applies to any given scan.

What do the grades mean?

GradeScoreReading
A90-100Excellent - highly likely to be cited
B80-89Good - likely to be cited
C70-79Average - several dimensions need work
D60-69Poor - significant GEO gaps
F0-59Critical - rarely cited by AI engines

What the scanner cannot tell you

A score is only useful if its limits are stated. These are the ones that matter:

  • It does not ask any AI engine about you. Nothing here measures whether ChatGPT currently recommends your brand. It measures whether your pages are in a state that makes being cited possible. Those are different questions, and conflating them is the most common overclaim in this category.
  • A refused request is not proof of a block. Large sites verify crawlers by IP address rather than user-agent, so a scanner’s request can be refused even when real GPTBot traffic is served normally. Whether a site blocks AI crawlers is answered here by robots.txt, which is the site’s stated intent, not by guessing from one response code.
  • Only the homepage is analysed. Site-wide problems on interior pages are invisible to a single-URL scan.
  • Performance checks are coarse. Real Core Web Vitals need a browser; this scanner deliberately does not run one. That is why the Delivery dimension carries the lowest weight.
  • Weights are a judgement, not a measurement. The dimension weights reflect published research and the fact that unreadable pages cannot be cited. They are stated in full so you can disagree with them specifically rather than vaguely.

The twelve dimensions

01AI Crawler Access

16%

If a crawler cannot read the page, nothing else on this list can matter. This is the only dimension where a failure makes the rest of the score moot.

Why this weight: Highest weight of the twelve. A page that is blocked or unreadable has no path to being cited, so it dominates the total.

  • 2 ptA request for /robots.txt returns HTTP 200 with a readable body.
    Not satisfied: No readable robots.txt at the domain root.
  • 6 ptParsing robots.txt into user-agent groups, none of gptbot, claudebot, perplexitybot, oai-searchbot or google-extended is disallowed from /. An exact agent group takes precedence over the * group, and the last matching rule wins.
    Not satisfied: At least one AI crawler is disallowed from the site root.
  • 6 ptThe inverse of the check above, reported by name so the blocked agents appear in the finding.
  • 1 ptrobots.txt contains a Sitemap: line.
  • 2 ptThe homepage returns HTTP 200 to a non-browser request that identifies itself honestly.
    Not satisfied: A non-200 homepage response. This is reported as its own finding and is not by itself proof that AI crawlers are blocked.
  • 2 ptThe homepage has no meta robots directive containing noindex.
  • 3 pt/sitemap.xml returns a body containing <urlset> or <sitemapindex>.
    Not satisfied: No sitemap, or a response that is not valid sitemap markup.

02Machine Readability

12%

JSON-LD is how a model binds your brand name, domain and product into one entity instead of inferring three unrelated strings.

Why this weight: Second highest. It is deterministic and fully in your control, and it is what makes entity disambiguation possible.

  • 5 ptAt least one <script type="application/ld+json"> block parses as JSON. Which blocks failed to parse is reported.
  • 4 ptParsed JSON-LD contains an @type of Organization, WebSite, Person or LocalBusiness.
  • 3 ptParsed JSON-LD contains an @type of FAQPage, Article, BlogPosting, HowTo, Product, SoftwareApplication or BreadcrumbList.

03Content Depth

11%

Generative engines select sources that answer a question thoroughly. Thin pages are rarely retrievable regardless of how well they are marked up.

Why this weight: High, because depth is a precondition for retrieval rather than a bonus on top of it.

  • 5 ptVisible text, after removing script, style and comment content, contains at least 800 word tokens. 300-799 is a partial pass.
  • 3 ptThe page contains at least 4 H2 elements. 2-3 is a partial pass.
  • 3 ptThe page contains both a <ul>/<ol> and a <table>. Having only one is a partial pass.

04Citability & Evidence

11%

Statistics, quotations and cited sources are the interventions with the largest measured effect in the published GEO research.

Why this weight: High and research-driven: this is the dimension most directly tied to measured visibility gains.

  • 4 ptBody text contains at least 5 numeric claims: percentages, currency amounts, multipliers of the form 3.2x, or bare numbers of three digits or more.
  • 3 ptThe page contains at least one <blockquote> element.
  • 3 ptThe page links to at least 2 external hosts on the authority list: arxiv.org, doi.org, nature.com, science.org, acm.org, ieee.org, springer.com, sciencedirect.com, any .gov or .edu host, wikipedia.org, github.com, developer.mozilla.org, web.dev, nih.gov, who.int or europa.eu. Own-domain links are excluded.
  • 1 ptThe page declares a rel=canonical link.

05Answer Readiness

10%

Answers are extracted as spans, not pages. A question-shaped heading followed by an immediate answer is the easiest thing for a retrieval system to lift intact.

Why this weight: Substantial, and largely mechanical to fix.

  • 4 ptAt least 3 H2 or H3 headings either end with a question mark or begin with how, what, why, which, when, who, where, is, are, do, does, can, should or will.
  • 3 ptParsed JSON-LD contains a FAQPage node. If instead the page uses <details>/<summary> markup without the schema, that is a partial pass.
    Not satisfied: No FAQPage schema. Note that this check reads structured data, not the word "faq" or "question" appearing anywhere in the HTML.
  • 3 ptOf the first 30 heading-then-paragraph pairs (H2/H3 optionally followed by wrapper elements then a <p>), at least 60% have an opening paragraph of 80 words or fewer.

06Trust & Authority

10%

E-E-A-T signals decide whether a model treats a claim as safe to repeat rather than something it should hedge.

Why this weight: Weighted equally with answer readiness: attribution is what separates a quote from a rumour.

  • 2 ptThe homepage was retrieved over https://.
  • 3 ptHrefs on the page include both an about-style path (about, company, team, who-we-are) and a contact-style path (contact, support) or a mailto: link.
  • 3 ptAuthorship is signalled by any of: a Person node in JSON-LD, an author property, rel="author", or a visible byline matching "by Firstname Lastname".
  • 2 ptJSON-LD contains a sameAs property.

07Semantic Structure

8%

Heading hierarchy and landmark elements are how a parser locates section boundaries at all.

Why this weight: Moderate. Mostly a one-time fix.

  • 3 ptThe page contains exactly one non-empty <h1>. Zero or several is a partial or failed result.
  • 3 ptAt least 3 of main, article, section, header, nav and footer appear as elements.
  • 2 ptThe raw server response contains at least 100 words of visible text without executing JavaScript.

08Metadata & Discoverability

7%

Title, description and canonical control what a search or answer surface can show about the page.

Why this weight: Moderate, and cheap to satisfy.

  • 3 pt<title> is between 15 and 65 characters long.
  • 2 ptThe meta description is between 50 and 160 characters long.
  • 1 ptAt least one meta property beginning with og: is present.
  • 1 ptThe <html> element carries a lang attribute.

09AI Context Files

5%

A cheap and optional signal. Google has stated it does not use llms.txt in Search, and crawler support is inconsistent.

Why this weight: Deliberately low. Most tools in this category weight llms.txt heavily and imply it is a ranking factor; the evidence does not support that, so neither does this score.

  • 4 pt/llms.txt returns more than 20 bytes and contains an H1, a > summary line and at least 3 markdown links. Missing any of those is a partial pass.
  • 1 ptrobots.txt is readable, so a crawler policy is published.

10Freshness

5%

Dated content is deprioritised in generated answers, and an undated page gives an engine nothing to reason about.

Why this weight: Low to moderate: a real but secondary signal.

  • 3 ptA machine-readable date is found in dateModified, datePublished or article:modified_time. Within 12 months is a full pass; older is a partial pass. If only the HTTP Last-Modified header carries a date, that is a partial pass.
  • 2 ptThe current calendar year appears in the visible body text.

11International Readiness

3%

Language and region markup decides which language market a page can be retrieved in at all.

Why this weight: Low for a single-market site, and the first thing to fix if you serve more than one language.

  • 2 ptAt least 2 distinct hreflang language codes are declared. Repeated attributes for the same language do not count, and x-default is a default marker rather than a language, so a single-language site declaring `en` plus `x-default` does not satisfy this.
  • 1 ptAn hreflang-style language tag is declared on the document.

12Delivery & Mobile

2%

A slow or non-mobile-readable page is dropped before any content analysis happens.

Why this weight: Lowest. The checks here are coarse on purpose; full Core Web Vitals need a real browser, which this scanner deliberately does not run.

  • 1 ptA meta viewport tag is present.
  • 1 ptThe uncompressed HTML response is under 500 KB.

Primary sources

Where this method makes a judgement about what generative engines favour, it follows published research rather than folklore. Read the sources directly:

This site holds itself to the same standard it measures: it publishes its own llms.txt, declares a single Organization entity in JSON-LD, and labels its own dates. You can verify all of it by scanning this domain.