mark-white

Scoring Methodology v1.0

How we score AI visibility: eight checks, published weights, a breakdown that sums to the headline number, and our own site's score.

Why this page exists

A score nobody can check is a claim, not a measurement.

We publish the checks, the weights and the limitations so any number we produce can be disputed on the merits. If a client or a competitor thinks a score is wrong, this page is where they show it.

The score

Technical readiness is a weighted score out of 100 across eight checks.

CheckMax pointsWhat it measures
robots.txt and AI crawler rules18Whether named AI crawlers are permitted, and whether citation bots are addressed explicitly
llms.txt18Presence, structure, sections, links, and a companion llms-full.txt
Schema.org structured data16Organization, WebSite, FAQPage and other types, and whether they resolve to one entity
Meta and Open Graph tags14Title, description, canonical, OG title, description and image
Content quality12Readable text volume, heading structure, depth
Freshness and language signals6Dates, language declaration, update signals
AI discovery files6agent-card.json, ai.txt and related discovery endpoints
Brand entity signals10Whether the site states machine-readably what the brand is, with external identifiers
Negative signalspenaltyPromotional density, keyword stuffing, boilerplate ratio, broken links

The breakdown sums to the headline number by design. Any score can be explained line by line, and the “where your missing points are” section of every report is generated from the gaps rather than written by hand.

Crawler access testing

Scoring the site is only half of it. We then test whether the crawlers can reach it.

Each homepage is requested as GPTBot, ClaudeBot and PerplexityBot, recording HTTP status, payload size and whether the response differs from what a browser receives. This is a real fetch. It is not inferred from robots.txt.

Two failure modes appear only through this test:

Refused at the edge. The robots.txt permits the crawler and the CDN answers 403. We have measured this on a market leader with the highest technical score in its field.

Delivered empty. The crawler receives HTTP 200 and a full payload containing no title, no headings and no body text, because the content is assembled by JavaScript in the browser.

Effective score is technical readiness weighted by crawler accessibility. A site no crawler can read cannot realise its readiness, so its effective score falls toward zero.

What the score does not measure

Stating this plainly matters more than the score itself.

It does not measure live citations. Technical readiness predicts whether a site can be read and cited. It does not count how often a brand appears in ChatGPT answers today. That requires a per-brand prompt set run against live engines, which is part of an engagement rather than the free report.

It does not measure recommendation. Being on a buyer’s shortlist is governed by web-wide consensus rather than by your own site. See the visibility ladder.

It is homepage-scoped in the free report. Full engagements score templates and key pages across the site.

It is a snapshot. Scores from different dates are not comparable. This is why our published studies republish the whole table rather than updating rows.

Limitations we know about

Single run. Published index scores come from one run per site on one date. No retries, no best-of.

Weights are a judgement. The eight weights reflect what we believe currently matters most. They are not derived from a controlled experiment, because no public dataset exists that would support one. When better evidence arrives, the weights change and the version number changes with them.

llms.txt carries 18 points and has no demonstrated effect on citation. We weight it because it is a strong proxy for whether a site has thought about machine readability at all, not because the file itself is a proven signal. That is a deliberate choice and a fair thing to argue with.

Crawler testing uses user-agent identification. A site could in principle treat our requests differently from real crawler traffic. We do not spoof IP ranges.

Why we measure repeatedly, and the research behind it

AI answers are probabilistic. The same question asked twice can return different sources, so a single observation is not a measurement.

Schulte, Bleeker and Kaufmann quantified this in Don’t Measure Once: Measuring Visibility in AI Search (arXiv 2604.07585). Their central finding: source overlap between consecutive days falls into the 34% to 42% range. Roughly two thirds of the sources cited for a query one day are different the next.

Two consequences follow, and both are built into how we work.

A one-off snapshot understates brand presence. The paper shows that any single reading of citation rate or mention rate can swing materially across repeated queries. A screenshot showing you are absent may simply be the run where you were absent.

Visibility is a distribution, not a number. We report a brand’s position across repeated runs rather than as a single figure, and the frozen prompt set exists so those runs are comparable to each other.

This is also why we distrust competitor claims built on a screenshot. Anyone can produce one run in which they rank first for their own category.

Reproducibility

Two different things are reproducible here and they should not be confused.

Site scoring is deterministic. The same site on the same day produces the same technical readiness score, because the measurement runs through a containerised pipeline rather than by hand.

Live citation measurement is not deterministic, and cannot be, for the reasons above. It is made comparable by holding the prompt set, the engines and the cadence fixed, then reporting the distribution across runs.

Any organisation we have published a score for can request their full report, including the raw audit output, and re-run it themselves.

Our own score

We run this methodology against our own site and publish the result.

Thysanos (thysanos.com).

Technical readiness: [score]/100.

Effective score: [score]/100.

Last measured: [date].

Open findings: [list them, including the ones not fixed yet.]

Publish this before launch and keep it current. A GEO company that has not scored its own site is the easiest objection a competitor will ever get, and publishing an imperfect number is more credible than publishing a perfect one.

Changelog

VersionDateChange
1.0[date]Initial published methodology

Every change to weights or checks gets a version bump and a row here. Historical studies keep the version they were measured under.

Swim in the spotlight.

Thysanos is the Generative Engine Optimization service of SyeniteLabs.

© 2026 · THESSALONIKI

SyeniteLabs SINGLE MEMBER P.C.
VAT: 803238622
ELGEMI.192798006000