Generative engine optimization is one of the few terms in this category that started somewhere respectable. Not an agency blog and not a product launch, but a paper with a benchmark and a reproducible result.
Most of what is written about it now has drifted a long way from that. This post covers what the original work actually claims, what holds up against what the engines themselves document, and which parts of the current GEO industry are selling something that does not exist.
Where the term comes from
GEO was formalised in a 2023 paper by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, later presented at KDD 2024. The authors define generative engines as systems that answer a query by synthesising information from multiple sources and summarising it, and they make the point that this leaves content creators with "little to no control over when and how their content is displayed."
Their contribution is two things. A framework for optimising content against those engines, and GEO-bench, a benchmark of queries across domains with the web sources needed to answer them, so that claims can be tested rather than asserted.
The headline result is that their methods boost visibility "by up to 40% in generative engine responses" on that benchmark.
Read that carefully, because it is quoted badly everywhere. It is an upper bound, on a specific benchmark, using the paper's own visibility metrics. It is a genuine finding. It is not a percentage any vendor can promise you on your category, and anyone repeating it as one has not read past the abstract.
What GEO shares with ordinary SEO
More than the acronym suggests.
Google's own guidance puts its generative features inside the existing system rather than beside it. The AI features documentation describes them as rooted in the core Search ranking and quality systems, which means being retrievable in ordinary Search is the precondition for appearing in an AI answer at all.
There is also a hard eligibility gate. Google's guidance is that nosnippet, data-nosnippet, max-snippet and noindex restrict how your content is featured in AI experiences. A directive left on a template silently disqualifies pages that are otherwise perfectly good.
And there is no markup shortcut. Google states that structured data is not required for generative AI search and there is no special schema.org markup you need to add. Keep schema for rich results. Do not buy it as a GEO product.
What is genuinely different
Three things, and they are the reason the term earns its existence.
The outcome is not a position. SEO optimises for a rank in a list. GEO optimises for inclusion in a synthesised answer, and that splits into two independent events: being named in the prose, and being cited as a source. A page can be cited while a competitor is named, which is its own problem with its own fix.
The retrieval set is wider than the query. Google generates related queries alongside the one typed and assembles from that broader pool, which it has described while expanding AI Mode. You end up competing for questions nobody asked, and a page that answers one narrow question thoroughly can surface for an adjacent one.
The answer is generated, not retrieved. This is the part that breaks conventional measurement. Two runs of the same prompt an hour apart can name different brands with nothing having changed. Everything about how you verify a GEO change follows from that, and most published case studies ignore it.
Engines are not interchangeable
A single "GEO score" assumes one target. There are at least six, and they behave differently.
OpenAI documents its crawlers separately by purpose, so blocking the training crawler and the search crawler are different decisions. Anthropic added web search to Claude in 2025, and in our measurements Claude cites fewer sources and is more willing to decline. Perplexity documents PerplexityBot and cites more heavily than anything else, which makes it the best early read on whether retrieval is working. Google runs AI Overviews and AI Mode as separate surfaces that disagree with each other on the same day.
Optimising for "AI" in the singular is how you end up strong on one surface and absent on another without knowing which. We report each engine on its own for that reason.
The honest list of what to do
Nothing here is exotic, which is itself informative.
- Be reachable. Audit robots.txt against each assistant's documented crawlers, and audit snippet directives. Binary, cheap, and it disqualifies more sites than anyone expects. While you are in there, consider adding an llms.txt file, which is cheap hygiene rather than the ranking lever some vendors sell it as.
- Answer questions in liftable prose. Assistants assemble from passages. One clear paragraph that answers a specific question beats the same information spread across a table and three footnotes.
- Be specific rather than promotional. Models lift concrete claims. "Six engines on every plan" is liftable. "Industry-leading AI visibility" is not.
- Fix how third parties describe you. The pages cited in your category are mostly not yours, so the roundups and comparison posts brief the model before your site does.
- Measure it properly. Which is the part everyone skips.
GEO, AEO and the acronym argument
Two terms circulate for roughly the same work, and the difference matters less than the volume of writing about it suggests.
The longer version of that argument is in what answer engine optimization is, including which parts of the AEO industry are being oversold.
Answer engine optimization grew out of the older world of featured snippets and direct answers, where the goal was to be the one box at the top. Its instincts are formatting instincts: clear question headings, concise definitions, structured answers.
Generative engine optimization assumes the answer is synthesised from several sources at once rather than lifted from one. Its instincts are about being one of the sources worth synthesising, and about the fact that being synthesised does not guarantee being named.
In practice the work overlaps almost completely. Be reachable, answer clearly, be described accurately elsewhere. If a vendor's pitch turns on which acronym they use, that is a signal about the vendor rather than about your strategy.
What does change between them is the unit of success. AEO asks whether you are the answer. GEO asks how often you are in the answer, which is a rate, and a rate needs a denominator. That is the point where most of this stops being a writing exercise and becomes a measurement one.
What the research does not say
It is worth being precise about the limits of the paper, since it is the only peer-reviewed thing most GEO content cites.
GEO-bench is a benchmark of queries with source documents. It is not a live measurement of ChatGPT, Gemini or AI Overviews as those products behave today, and the paper does not claim it is. Generative engines have changed substantially since 2023, including Google shipping AI Mode as a separate surface with its own retrieval.
So the honest reading is that the paper establishes the problem is real, tractable and measurable, and that content changes can move visibility meaningfully on a controlled benchmark. It does not establish which specific tactic will work on your category on a given engine this quarter. Only running the prompts will tell you that, which is the unglamorous conclusion this whole field keeps arriving at.
Why measurement is the whole game
If the answer is generated fresh each time, a before-and-after screenshot is worthless. So is a single check, in either direction.
What produces evidence is a fixed prompt set, run on a schedule, with the wording held constant, on each engine separately, with every answer stored in full alongside its citations. Every rate carries the count behind it. Below five checks you show a fraction, because "2 of 4" is honest and "50%" implies a precision four runs cannot support. That is the substance of our methodology and it is deliberately boring.
Thirty or so runs per prompt per engine is roughly where a real change separates itself from ordinary variance. Changing the wording of a tracked prompt restarts the series, because a trend measured across two different prompts is not a trend.
For the definition sitting underneath all of this, see what AI visibility is. For how the assistants differ in practice, how AI assistants decide who to name. And if you want the measurement without building it, every plan here runs all six engines and keeps every answer.
GEO is a real discipline with a real paper behind it. Most of what is sold under the name is the measurement problem wearing a costume.

