mark-white

What Is GEO? Generative Engine Optimization Explained (2026)

Generative Engine Optimization is the practice of making a site readable, extractable and citable by AI assistants. What GEO is, what the research shows, and what it is not.

What is Generative Engine Optimization?

Generative Engine Optimization (GEO) is the practice of structuring a website so AI assistants can read it, extract passages from it, and cite it as a source in their answers. It covers three things: whether AI crawlers can reach your pages at all, whether your content is machine-readable once they arrive, and whether your brand is described consistently enough across the web for a model to trust it.

The term comes from research. In 2024, a team from Princeton University and Georgia Tech published GEO: Generative Engine Optimization at KDD 2024, introducing GEO-bench, a benchmark of 10,000 queries across eight domains. They tested nine optimization methods and measured which ones changed whether a source appeared in a generated answer.

That paper is why the discipline has a name, and it is still the only large public study that tested the methods rather than asserting them.

What the research actually found

Across roughly 10,000 queries, the Princeton and Georgia Tech team measured visibility improvements of 22% to 41% depending on method and domain.

Three methods did most of the work:

MethodEffectWhat it means in practice
Cite sourcesstrongest groupLink to authoritative references inside the content
Add statisticsstrongest groupSpecific numbers, attributed, dated
Add quotationsstrongest groupNamed experts with title and organisation
Authoritative tonemoderateWrite with demonstrated expertise
Improve claritymoderateSimplify without losing precision
Keyword stuffingnegativeReduces visibility. It is not neutral, it costs you.

The last row is the one most vendors skip. Density tactics carried over from 2010s SEO measurably lower AI visibility. Writing more keywords into a page makes it less likely to be cited, not more.

The short version: a page that cites its sources and shows its numbers gets picked up. A page that repeats its keywords gets passed over.

Why measurement has to be continuous

One number from the follow-up research shapes how this work has to be run.

Schulte, Bleeker and Kaufmann measured how much AI answers move between runs. In Don’t Measure Once (arXiv 2604.07585, 2026) they found source overlap between consecutive days falls in the 34% to 42% range. Roughly two thirds of the sources cited for a query one day are different the next.

Two things follow. A single screenshot proves nothing, in either direction. And visibility is a distribution rather than a fixed position, which means it has to be measured on a fixed prompt set at a fixed cadence to mean anything at all. That is why our engagements are monthly rather than one-off, and why the prompt set is frozen on day one.

Why this matters now

The shortlist your buyers consider is being written inside AI assistants, and it closes before they contact anyone.

AI is now the single biggest influence on B2B shortlists. G2’s 2025 Buyer Behavior Report, covering 1,100 B2B decision-makers globally, ranked generative AI chatbots the number one source shaping vendor shortlists at 17.1%, ahead of software review sites at 15.1%, vendor websites at 12.8% and peer recommendations at 8.9%. Machine Relations puts LLM use in vendor research at 94% of B2B buyers, and G2’s March 2026 follow-up found 51% now start their research in an AI chatbot more often than in Google.

92% of buyers say AI shaped their shortlist and 83% say it influenced the final decision, in a Semrush survey of 622 US business professionals fielded in March and April 2026. Both figures describe the point at which a deal is won or lost rather than anything at the awareness stage.

And the shortlist closes before anyone picks up a phone. 6sense surveyed around 4,000 buyers and found the winning vendor was already on the buyer’s day-one shortlist 95% of the time. Four deals in five go to the vendor the buyer privately favoured before making contact, and 94% of buying groups rank that shortlist before speaking to a single seller. Eighth place on a results page still put you in front of somebody. Eighth place against an answer naming three companies reaches nobody at all.

Which cuts both ways, and that is the opportunity. In G2’s 2026 survey, 69% of buyers chose a different vendor than they had planned after researching with a chatbot, and a third bought from a company they had never heard of before the assistant named it. Brand recognition, meanwhile, is why only 7% of buyers notice a vendor in an AI answer at all. What earns the slot is a clear description that matches the use case, cited at 53% in the same Semrush survey.

The click is disappearing from the searches AI touches.  racked 68,879 real Google searches by 900 US adults and found users clicked a esult on 8% of searches when an AI summary was present, against 15% when it as not. Pew Research: About 1% clicked a source cited inside the summary. Sixty percent f searches beginning with who, what, when or why returned one. Similarweb puts overall zero-click Google searches at 68%.

How often an AI Overview appears depends on who is counting, and every count is rising. Conductor measured 25.11% across 21.9 million queries in late 2025. BrightEdge measured 48% across nine commercial verticals in February 2026, up 58% in a year. The gap is the keyword panel each study used, not a disagreement about direction.

AI referral traffic is growing faster than any channel in recent memory. Adobe Analytics recorded traffic to US retail sites from generative AI rising 1,200% between July 2024 and February 2025. Across 13,770 domains and 3.3 billion sessions, Conductor puts AI referrals at 1.08% of all website traffic, with 87.4% of it coming from ChatGPT.

ChatGPT passed 900 million weekly users in February 2026, up from 800 million four months earlier, and crossed a billion monthly app users in June.

And these visitors arrive closer to a decision than any other channel delivers them. Semrush found the average AI search visitor 4.4 times as valuable as the average organic one, measured across 500+ topics. At the top of the range, Ahrefs reported that on its own site AI search was 0.5% of traffic and 12.1% of signups, a 23x conversion rate. In one Seer Interactive client case study, ChatGPT traffic converted at 15.9% and Perplexity at 10.5%, against 1.76% for Google organic.

Adobe found the same pattern in engagement. Visitors arriving at US retail sites from generative AI browsed 12% more pages, bounced 23% less, and showed 8% higher engagement than visitors from every other source.

Someone who arrives from an AI answer has already had their question answered and their options narrowed, so they reach your site closer to a decision than any other channel delivers them.

Why the next twelve months decide it

Three things are happening at once, and they push in the same direction.

  • The shelf has three slots and it is being stocked now. A search results page offered ten links and a scroll, where an AI answer names two or three companies and stops. Since 95% of deals go to a vendor already on the buyer’s day-one shortlist, an answer that names three companies leaves no gradual middle to climb through.
  • Assistants reinforce what they already cite. Every answer that names a company makes the next answer more likely to name it again, so the brands establishing themselves in a category now are becoming the default a model reaches for. Displacement gets harder every month that passes.
  • And almost nobody has claimed their category yet. When we measured sixteen major European tour operators in July 2026, the highest score in the entire field was 68 out of 100. The average was 44. Fourteen of the sixteen had no llms.txt at all, and four were effectively invisible to AI crawlers without knowing it. That is a market-leading position sitting unclaimed, in an industry turning over billions.

We find the same picture in most sectors we measure. The work required to lead a category today is smaller than it will ever be again, because the bar is still on the floor and the people who set it will be the ones the assistants learn first.

In eighteen months this stops being a land grab and becomes a displacement problem, and displacing an incumbent costs several times what establishing a position does today.

The three things GEO actually covers

1. Access: can AI crawlers reach you?

Each AI platform sends a differently named crawler. GPTBot and ChatGPT-User for OpenAI, ClaudeBot for Anthropic, PerplexityBot for Perplexity, Google-Extended for Gemini and AI Overviews, Bingbot for Copilot. Blocking one means that platform cannot cite you, regardless of how good your content is.
Two failure modes are common and neither is visible from a robots.txt file:

  • A hard block at the CDN. The robots.txt invites the crawler, and Cloudflare or Akamai answers it with a 403. We have measured sites with the strongest technical content in their market scoring zero effective visibility for exactly this reason.
  • A page that renders to nothing. The crawler gets HTTP 200 and the full payload, and the payload is a JavaScript shell with no title tag and no body text. The content exists only after JavaScript runs, and most AI crawlers do not run it.

Testing this needs a real fetch as each named agent. Reading the robots.txt file tells you nothing about either failure.

2. Structure: can they use what they find?

AI systems extract passages, not pages. Content that gets cited shares a shape:

  • A direct answer in the first 40 to 60 words under each heading
  • Headings phrased the way people actually ask
  • Tables for comparisons, numbered lists for processes
  • Statistics with a source and a date attached in the sentence
  • Structured data in JSON-LD saying what the page is and who published it
  • A visible last-updated date and a named author

3. Entity: does the web agree on who you are?

A model needs to resolve your brand to a single, stable thing. That comes from Organization schema with a consistent @id, sameAs links to Wikidata, Wikipedia, LinkedIn and Crunchbase, and consistent descriptions of the business across every third-party surface where it appears.

This is the part most site-only work misses. Brands are cited from third-party sources far more often than from their own domain, so an entity that only exists on its own website is an entity a model is unsure about.

Should you block AI crawlers?

This deserves a straight answer, because the trade is real and most GEO vendors pretend it is not.

AI crawlers currently take far more than they return. Cloudflare Radar
measured crawl-to-referral ratios on 31 May 2026:

CrawlerPages crawled per referral sent back
ClaudeBot10,300 : 1
GPTBot903.8 : 1
PerplexityBot192.9 : 1
Googlebot5.2 : 1

That is why GPTBot is the most-blocked AI crawler on the web, appearing in 5.52% of DISALLOW rules across Cloudflare’s network in Q1 2026, and why the New York Times, Reuters, the Guardian and Bloomberg all block it.

If you sell advertising against pageviews, blocking is a defensible business decision. If you sell a product or a service and you want to be on the shortlist when someone asks an assistant for a recommendation, blocking removes you from consideration entirely.

The middle path most businesses want: allow the search-and-cite crawlers, block the training-only ones such as CCBot, and keep price-scraping traffic out with the controls that were always doing that job. Every crawler identifies itself, so this is a configuration decision rather than an all-or-nothing switch.

What GEO is not

  • It is not a rebrand of SEO. Good technical SEO is the foundation, and GEO adds extraction structure and entity work on top. Sites that skip the fundamentals do not get to skip them.
  • It is not separate content written for machines. Google’s AI features guidance
    is explicit that writing AI-targeted content variants and breaking pages into fragments for AI risks the scaled content abuse policy. Write for the reader, organise so a passage lifts cleanly.
  • It is not llms.txt. The file is worth shipping and we ship it. It is also sitting on about 10.13% of domains as of 2026 with no demonstrated correlation to citation rate, Google has said publicly it does not support it and has no plans to, and no major model provider has committed to using it as a signal. Anyone selling llms.txt as the answer is selling a text file.
  • It is not a guarantee. Nobody controls what a model outputs. What can be guaranteed is measurement, implementation, and a re-runnable baseline that shows whether the work moved anything.

Frequently Asked Questions

Sources

Swim in the spotlight.

Thysanos is the Generative Engine Optimization service of SyeniteLabs.

© 2026 · THESSALONIKI

SyeniteLabs SINGLE MEMBER P.C.
VAT: 803238622
ELGEMI.192798006000