Generative Engine Optimisation (GEO)
Last updated
Generative Engine Optimisation (GEO) is the practice of structuring and positioning content to be retrieved, parsed, and cited by AI-powered answer engines. It overlaps with traditional SEO heavily, but the goals diverge in places: GEO optimises for citation, not click-through.
Google’s stance: GEO is SEO
Google’s May 2026 AI optimisation guide states its own position plainly:1 “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.” From Google’s perspective, there is no separate GEO discipline. Sites that perform well in traditional search, through quality content, clear structure, and authoritative authorship, perform well in Google’s AI-generated answers without special optimisation. The guide also has a mythbusting section listing five things you can ignore for Google Search: llms.txt and other special markup, content chunking, rewriting content for AI systems, seeking inauthentic mentions, and overfocusing on structured data. Other AI search platforms (Bing, Perplexity, ChatGPT Search) have not issued equivalent guidance and may weight signals differently.
How does citation-focused optimisation differ from ranking-focused optimisation?
Traditional SEO optimises a page to rank in a list of results. The user clicks the result, lands on the page, reads the content, and possibly converts. The page is the destination.
GEO optimises a page to be retrieved by a generative system, which then synthesises an answer from one or more sources and presents that answer (with citations) inside its own interface. The page is no longer the destination, it is a source for synthesised answers.
The implications:
- Click-through rate matters less; citation rate matters more. A page cited in AI Overviews 10,000 times that earns 200 clicks is performing differently than one ranked #3 organically with the same impression volume. Both have value, but the value is different in kind.2 The compensating factor comes from Knotch’s own client-panel tracking, quoted inside Conductor’s benchmarks report rather than measured by it: LLM-referred visitors converted at twice the rate of other traffic sources, in a third of the number of sessions.3
- Brand visibility extends beyond the click. When a page is cited, the brand is named in the response even if the user never visits. That visibility has compounding effects on entity recognition.
- Passages, not pages, are the unit of retrieval. AI engines extract specific passages that answer specific questions. A 4,000-word article rarely gets cited as a whole; one paragraph from it might. Microsoft’s Web IQ, a Bing-based grounding API launched in June 2026, evaluates passages against GDSAT (completeness, freshness, authority) independently of the page they sit on.4 Microsoft says Web IQ grounds Copilot directly and supplies ChatGPT with “some of its web answers”,5 so passage-level selection governs a substantial share of what those two surfaces cite.
What does the evidence actually say works?
Most GEO advice is inference from how retrieval systems appear to behave. Two bodies of evidence are not, and they point in different directions: the paper that coined the term, and the benchmark that later tested the whole class of methods it started.
The paper that coined the term tested content edits against generative engines and measured the change in visibility.6 Three of its interventions stand out, and they are the same three:
- Add statistics. Replacing vague quantities with specific figures.
- Add quotations. Quoting named, credible sources directly.
- Add citations. Attributing claims to their sources.
The strongest methods improved on the baseline by 41% on Position-Adjusted Word Count and 28% on Subjective Impression, the study’s two visibility measures, with reported gains of up to 40% overall.6 The paper also found the effect is not uniform: statistics helped most in domains like law and government and on opinion-shaped questions, while quotation addition did most work in explanatory and historical content.6
Two caveats keep this honest. The first is scope. GEO-bench supplies the engine with “relevant web sources to answer these queries” and the baseline measures “the impression metric of unmodified website sources”, so what the paper optimises is a source’s share of an answer it was already being shown for. It is evidence about winning more of an answer, not about being retrieved into one in the first place. The second is age: the study predates the current generation of AI search products, so treat it as directional evidence about how generative systems weigh evidential content rather than a live specification of any one engine.
The later test is less encouraging, and it is the better controlled of the two. C-SEO Bench, accepted at NeurIPS 2025, is the first benchmark to evaluate this class of methods across multiple tasks, multiple domains and multiple adopters, the dimension earlier tests left out: two search tasks, question answering and product recommendation, with three domains each, plus an evaluation protocol that varies how many competing documents apply the same technique at once. Its finding is that “most current C-SEO methods are not only largely ineffective but also frequently have a negative impact on document ranking, which is opposite to what is expected”. Traditional SEO aimed at improving a source’s rank within the LLM context was “significantly more effective”. Gains also shrank as adoption rose, which the authors describe as a “congested and zero-sum” dynamic.7
Read together, the two produce a narrower conclusion than either headline. The interventions the original paper tested, stating things precisely, quoting named sources and attributing claims, are also the ones that make a page better on ordinary quality grounds, which is why they survive a test most of the category fails, and why they converge with what Google says anyway. Techniques that exist only to game a generative engine are the ones C-SEO Bench finds ineffective or actively harmful, and they degrade further as more sites adopt them. The evidence supports writing more precisely. It does not support adopting a technique because it is described as a GEO tactic.
GEO, AEO, LLMO: why there are so many names
The practice described on this page goes by several names. GEO is used here because it has the clearest academic provenance: it was coined in a 2023 research paper by Aggarwal and colleagues examining how content performs in generative answer engines.6 The others are in common use.
AEO (Answer Engine Optimisation) is the oldest term. It predates generative AI, originally describing optimisation for voice assistants and featured snippets: surfaces that return a direct answer rather than a list of links. Since large language models became prominent in search, AEO has been used more broadly and often interchangeably with GEO. Some practitioners draw a narrower distinction: AEO for Google’s own AI features (AI Overviews, AI Mode), GEO for third-party systems such as Perplexity and ChatGPT. Neither usage is wrong; the overlap is near-total.
LLMO (Large Language Model Optimisation) emphasises the model layer: influencing what LLMs incorporate into their training data, beyond what they retrieve in real time. In practice, the content tactics for LLMO are indistinguishable from GEO. The framing differs; the work does not.
AIO appears in two contexts: as shorthand for AI Overviews (the Google feature), and as a generic label for any AI-related optimisation. The ambiguity makes it imprecise as a standalone term.
AI SEO is a catch-all. It covers the application of SEO thinking to AI search surfaces, which is precisely what GEO and AEO already describe.
None of these represent separate disciplines with separate playbooks. The content signals that earn citations are consistent across all of them: accuracy, clear structure, direct answers, credible authorship, cited sources. The label matters less than understanding why AI retrieval systems favour certain content patterns.
What does GEO content look like?
The patterns that correlate with citation across multiple AI surfaces (AI Overviews, Perplexity, ChatGPT Search) are consistent.
Direct, definitional opening sentences. The first sentence of a section should answer the question that section addresses. Burying the answer in a third paragraph reduces the probability that a retrieval system extracts it cleanly, and this one is measured rather than inferred: benchmarking across the retrieval pipeline found the embedding models that do the initial passage matching lost an average of 15.6% of retrieval performance when the relevant information sat later in the passage, while keyword matching and rerankers were largely unaffected.8 The effect therefore bites at the stage that decides which passages are considered at all.
Question-shaped headings. H2s and H3s phrased as questions match conversational queries closely. “What is X?” matches the wording someone actually uses more closely than “Understanding X” does, even though the two cover the same material. No published study measures the size of that difference, so treat it as reasoning about how retrieval matches text rather than a measured gain.
Stand-alone passages. Each section should make sense extracted on its own. Heavy use of pronouns referring back to earlier sections (“As we saw above…”) makes a passage harder to use as a citation.
Explicit attribution and citation. AI systems favour content that itself cites primary sources. Linking out to original research, official documentation, and named experts is a trust signal.
Structured data formats. Tables, lists, definition pairs, and step-by-step instructions are easier to parse than dense prose. They also map cleanly onto the formats AI engines tend to render in their answers.
Named author with verifiable credentials. Anonymous content from a faceless brand competes from a weaker position than content with a named expert author whose authority can be cross-referenced.
Neutral framing over self-promotion. AI answers reward content that reads as an even-handed authority, not a sales page. A June 2026 analysis of 100 B2B “best software” queries by Lily Ray found Google’s AI Overviews frequently cited a brand’s own “best [category]” listicle as a source while leaving that brand out of the recommendation, around 69% of the time, recommending competitors or Reddit, Forbes, and YouTube instead.9 Being cited as a source and being the recommended answer are different outcomes; self-serving framing earns the first and loses the second.
Is GEO only about on-page content?
No, and treating it that way is the most common mistake in the discipline. Everything above concerns the page itself: its structure, its passages, its attribution. That work determines whether a page can be cited cleanly once a retrieval system reaches it. It does not determine whether the brand is a candidate in the first place.
Ahrefs’ December 2025 analysis of 75,000 brands found the strongest off-site correlates of AI visibility were not links but mentions. YouTube mentions led, at 0.735 on ChatGPT, 0.736 on AI Mode and 0.740 on AI Overviews, with branded web mentions behind them at 0.656, 0.709 and 0.664, and a weak 0.19 to 0.24 for backlinks.10 On the Spearman scale those run from 1, the two rankings matching exactly, to 0, no relationship at all. The authors are careful to note that correlation is not causation, and the confound is obvious enough that it should temper any strong claim: brands that get written about are usually brands that are already established, and establishment drives both variables. An earlier Ahrefs study found the same correlation was strong for AI Overviews but weak for Perplexity and ChatGPT, so the effect is unlikely to be uniform across platforms.11
Set the correlation aside and a structural argument remains. Each engine draws from its own source pool, and those pools diverge: Semrush found in July 2025 that Google’s AI Mode shared only around 54% of its cited domains with Google’s own organic top 10, and that ChatGPT’s citations tracked Bing more closely than Google.12 That window predates both Gemini 3 and the August 2026 change to ChatGPT’s fan-out, so treat the divergence as the durable finding and the specific overlap figure as of its date. A page you own, however well structured, sits in one index. Coverage across third-party publications reaches source pools your domain cannot, which is why digital PR and brand mentions sit inside a GEO programme and not beside it. The on-page work makes a page citable; the off-page work makes the brand a candidate for citation.
What GEO is not
GEO is not “writing for robots” in the keyword-stuffing sense. AI retrieval models are evaluating the same quality signals that Google’s quality raters look for, just programmatically. Pages that read as machine-targeted (repetitive phrasing, keyword density above natural language norms, content padding) are penalised by both human readers and the systems learning from them.
GEO is also not a wholesale replacement for traditional SEO. The two share most of their underlying signals, which is Google’s own position on its systems: a site optimised well for traditional SEO is already most of the way to being well-optimised for AI retrieval. The part that does not overlap, retrieval and citation rather than ranking and clicks, is where the discipline lives.
GEO is also not a strategy for social discovery. Social search on platforms like YouTube, Reddit, and Instagram has distinct signals from traditional search. YouTube SEO, Reddit SEO, and Instagram search each require platform-specific approaches, covered under search everywhere optimisation.
Tactics that do not work
Google’s May 2026 guide has a mythbusting section, “what you don’t need to do”, naming five things you can ignore for Google Search.1 Those five are marked below. The last two items are not Google’s claims: they are consistent with its guidance but stated here on other grounds.
llms.txt files and other special markup. (Google’s guide.) Creating machine-readable files is unnecessary for Google’s AI systems, which parse web pages directly. Google adds that maintaining one “will neither harm nor help” visibility, because Search ignores it. Do not publish expecting SEO or GEO benefits in Google. Anthropic publishes its own llms.txt for documentation purposes and co-developed the llms-full.txt format with Mintlify; whether Claude reads other sites’ llms.txt at inference time is unconfirmed.13
Content chunking. (Google’s guide.) There is no requirement to break content into small pieces for AI to understand it, and Google states there is no ideal page length. Artificially fragmenting paragraphs to trigger retrieval signals poor editorial quality and reads unnaturally.
Rewriting content just for AI systems. (Google’s guide.) Google’s systems understand synonyms and intent, so there is no need to capture every phrasing a user might try. Content rewritten to sound machine-targeted underperforms naturally written content.
Seeking inauthentic mentions. (Google’s guide.) Google’s generative features do surface what is said about a brand across the web, but it names manufactured mentions specifically as less useful than they appear, because core ranking depends on content quality and separate systems block spam. This is the boundary on the off-page argument above: coverage earned because the work is worth covering counts, and coverage manufactured to be counted does not.
Overfocusing on structured data. (Google’s guide.) Structured data is not required for generative AI search and there is no special schema.org markup to add, though Google recommends continuing to use it for rich result eligibility. The separate failure mode is schema that misdescribes the page: marking up content as FAQ or HowTo when there is no genuine FAQ or procedure adds noise rather than eligibility.
Over-structuring content. Not named by Google. Breaking every paragraph into lists or forcing every heading into a question produces poor readability, and Google’s own framing throughout the guide is that structure should serve human readers. Reshape a heading where the section genuinely answers that question, and leave it alone where it does not.
Volume without substance. Google addresses this outside the mythbusting list, in its content guidance: publishing at scale to capture more surface “violates Google’s scaled content abuse spam policy”, and “a high quantity of pages doesn’t make a website higher quality or more relevant to users”.1 Scale without human review optimises against the signals that earn citations.
The common thread: tactics treating GEO as a separate game to be won independently of content quality will underperform. The sites earning consistent AI citations are those that would have ranked well in traditional search anyway. Understanding how AI search works, particularly the citation quality filters applied before a passage is included in an answer, explains why AI content risks, including hallucination and E-E-A-T erosion, result in citation exclusion rather than just ranking drops.
How do you start with GEO?
- Audit existing content for retrievability. For your top 20 pages, check whether each major section opens with a clear answer to a specific question. If not, rewrite the opening.
- Reshape headings into question form where the section content answers a question. Keep the rewrites natural; don’t force every heading into a question if it isn’t one.
- Add schema markup that describes what the page actually is. Article, Product, and Organization types cover most pages and help retrieval systems resolve what a page covers and who published it. FAQPage and HowTo stopped producing rich results in Google Search in May 2026, though the markup still helps AI systems extract question-and-answer pairs from pages that genuinely contain them. Adding either to a page with no real FAQ or procedure is the schema stuffing described above.
- Strengthen author attribution. Visible bylines, author bio boxes, Person schema with
sameAslinks to LinkedIn and other authoritative profiles. - Build off-page coverage deliberately. Audit where the brand is mentioned across publications the engines actually cite, then close the gaps through digital PR, original research, and expert commentary. This is the slowest of the five steps and the one most often skipped, because it sits outside the SEO team’s direct control.
- Track citation alongside rankings. In GSC, monitor CTR trends against stable impressions in the Performance report (Web search type): a falling CTR is the primary signal that an AI Overview is absorbing clicks. There is no dedicated AI Overview filter. Monitor AI-driven traffic in GA4’s AI Assistant channel, which captures ChatGPT, Claude, Gemini and Perplexity automatically.14 Supplement with manual sampling across Perplexity, ChatGPT, and other surfaces for citation tracking where referrer data is incomplete.
The GEO question for any piece of content
When deciding whether to publish, expand, or rewrite a page, ask: would a retrieval system extracting a passage from this page produce something a user would find useful and trustworthy? If yes, the page is GEO-ready. If the best passage requires several paragraphs of context to make sense, restructure until a stand-alone passage exists.
Frequently asked questions
What is the difference between GEO, AEO, LLMO, and AI SEO?
They largely describe the same practice under different names. AEO (Answer Engine Optimisation) is the oldest, originating in the voice search and featured snippet era. GEO (Generative Engine Optimisation) is more precise, coined in academic research to describe optimisation for systems that synthesise answers from retrieved content. LLMO (Large Language Model Optimisation) emphasises influencing LLM knowledge rather than real-time retrieval; in practice the tactics are nearly identical to GEO. AIO and AI SEO are broader terms with no distinct practice behind them. The content signals that earn citations are consistent across all these surfaces.
Is GEO a real discipline or just SEO rebranded?
Both. The underlying signals are mostly shared with traditional SEO. The framing, measurement, and content patterns are different enough to warrant separate language. Treat GEO as the part of SEO concerned with retrieval and citation in generative systems, not as a replacement.
Can I optimise for one AI engine without affecting others?
The content work transfers; the source pools do not. Writing that is accurate, well structured and clearly attributed helps on every surface, so there is little reason to produce engine-specific versions of a page. What does not transfer is which sources each engine draws on: they overlap only partially, and the off-page correlates differ by platform, as the sections above set out. Expect the on-page work to carry across and the visibility it produces to vary by engine.
Does GEO require sacrificing traditional SEO performance?
Rarely. Most GEO improvements (clearer structure, direct answers, better schema, stronger attribution) also help traditional rankings. Cases where the two diverge meaningfully are uncommon.
Footnotes
-
Optimizing your website for generative AI features on Google Search — Google Search Central ↩ ↩2 ↩3
-
Microsoft Web IQ (Microsoft names “AI platforms like Copilot and OpenAI”), with Jordi Ribas, President of Search and AI, quoted in Microsoft releases Web IQ, powered by Bing but designed for how AI-agents search — Search Engine Land ↩
-
GEO: Generative Engine Optimization — Aggarwal et al., arXiv 2311.09735 ↩ ↩2 ↩3 ↩4
-
C-SEO Bench: Does Conversational SEO Work? — Puerto et al., arXiv 2506.11097. Accepted at NeurIPS Datasets & Benchmarks 2025. ↩
-
An Empirical Study of Position Bias in Modern Information Retrieval — arXiv 2505.13950 (EMNLP 2025 Findings) ↩
-
Google AI Overviews cite self-serving listicles, but recommend competitors 69% of the time — Search Engine Land ↩
-
Top Brand Visibility Factors in ChatGPT, AI Mode, and AI Overviews (75k Brands Studied) — Ahrefs. Spearman correlations, December 2025. The authors explicitly caution that correlation is not causation. ↩
-
Google Seems More Biased Towards Big Brands Than ChatGPT and Perplexity — Ahrefs. Based on ~76.7M AI Overviews and roughly 950,000 prompts each for ChatGPT and Perplexity, June 2025. ↩
-
How Google’s AI Mode Compares to Traditional Search and Other LLMs — Semrush. 5,000 keywords, 150,000+ unique citations, July 2025. ↩
-
Google Analytics AI Assistant traffic measurement roll-out — Google Analytics LinkedIn ↩