← All Articles Radar Editorial
Market Trends Deep Dive

Generative Engine Optimization Has Actual Research Behind It. Most of What Gets Sold as GEO Doesn't

By AI SaaS Radar Team · Sep 2026 · 7 min read

Generative Engine Optimization is not a term a marketing agency invented. It comes from a peer-reviewed paper: "GEO: Generative Engine Optimization," by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, published at KDD '24 in August 2024. That matters, because the paper ran controlled experiments, and the results are considerably narrower and more specific than the advice now being sold under the same three letters.

What the paper actually tested

The researchers built a benchmark of queries, generated answers with a generative engine, and then systematically modified source websites to see which modifications increased that source's visibility in the generated answer. Crucially, they did not measure "did the page rank." They measured how much of the generated answer was attributable to the page, using metrics like position-adjusted word count and subjective impression.

That distinction is the whole discipline in one sentence. In classic search, the unit of success is a ranked link. In a generative answer, the unit of success is how much of the answer came from you, and how prominently. A page can be retrieved, contribute one clause buried in paragraph four, and be functionally invisible to the person reading. The paper's headline result, visibility improvements of up to 40%, is a statement about that, not about rankings.

What worked

The two standout methods were Statistics Addition and Quotation Addition: rewriting source content to include relevant quantitative data, and to include direct quotations from credible sources. Citing sources also performed well. The best-performing methods improved on the baseline by 41% on position-adjusted word count and 28% on subjective impression.

There is a coherent reason for this, and it is not mystical. A generative engine assembling an answer is looking for material it can lift and attribute with low risk. A sentence containing a specific number, or a quotable line with a named source attached, is high-value raw material: it can be dropped into an answer as a concrete claim. A paragraph of unsourced adjectives about being a leading provider of innovative solutions is not liftable. It says nothing a model can safely repeat.

What did not work

The methods that mirror classic SEO instincts performed worst. Keyword stuffing, the reflex move of cramming the target query into the page, did not reliably improve visibility, and in the paper's results was among the weakest interventions tested. That is the single most useful finding for anyone evaluating a GEO vendor, because a substantial amount of what is marketed as GEO in 2026 is keyword-density tooling with a new label on the box.

Why the research is necessary but not sufficient

Three caveats keep this honest. First, the paper predates the current generation of engines by roughly two years, and the retrieval stacks behind ChatGPT, Gemini, Claude, Perplexity and Google's AI Mode have all changed substantially since. The direction of the findings has held up better than the magnitudes.

Second, the experiments optimize a source page given that it was already retrieved. Getting into the candidate set in the first place is a separate problem, governed largely by conventional authority, crawlability and whether the engine's index has you at all. GEO does not replace that work; it sits on top of it.

Third, the results are averages across a benchmark. Your category may behave differently, and the only way to know is to measure your own prompt set rather than inherit someone else's conclusions.

The practical read

If you want a filter for GEO advice, use this one: does the recommendation make your content more liftable by a machine assembling an answer, or does it just make the page look more optimized to a human auditor? Specific numbers with dates and sources, direct quotes with attribution, clear declarative answers to actual questions, and comparison tables are liftable. Keyword density, word-count minimums and semantically-related-term checklists are not, and the research says so explicitly.

Stay ahead of the AI SaaS market

Sourced, dated analysis on security, funding, and benchmarks. Straight to your inbox.

No spam. Unsubscribe anytime.