Blog · Findings of the original Generative Engine Optimization (GEO) research paper

What the Original GEO Research Paper Actually Found

Priya Bothra · October 2, 2026

The original GEO research paper found that rewriting a web page to add quotations, statistics and citations to credible sources made that page noticeably more visible inside AI-generated answers, with the best single method improving one visibility metric by about 41% over an unoptimized baseline. It also found that keyword stuffing, the classic SEO shortcut, did little or even hurt, and that lower-ranked pages gained far more from these edits than pages already ranked first.

Those headline findings get repeated constantly, often stripped of the context that makes them useful. This article explains what the paper tested, how it measured visibility, which results held up on a real AI search product, what the study did not show, and how to apply it without overreading a single experiment.

What is the original GEO paper?

The original GEO paper is "GEO: Generative Engine Optimization" by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, with authors from Princeton University and the Indian Institute of Technology Delhi. It was first posted to arXiv in November 2023 and published at KDD 2024, the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, held in Barcelona in August 2024.

The paper did two things that shaped the industry vocabulary. It named the practice "generative engine optimization," defined as helping content creators improve how visible their content is in responses from generative engines. It also introduced GEO-bench, a benchmark designed to test that visibility in a repeatable way.

The paper treats the AI system as a black box. The authors did not have access to model internals. They changed the text of source pages and observed how much of the generated answer drew on those pages.

How was the GEO study set up?

The GEO study simulated an AI search engine in two steps: fetch sources, then generate an answer from them. According to the full paper, the test engine pulled the top 5 Google results for each query and passed them to GPT-3.5-turbo, which wrote an answer with inline citations. The team sampled 5 responses per query at a temperature of 0.7 to reduce random variation.

What is GEO-bench?

GEO-bench is a set of 10,000 queries assembled from nine sources, including MS MARCO, ORCAS-1, Natural Questions, ELI5, LIMA, AllSouls essay questions, Davinci-Debate, Perplexity.ai Discover and queries generated with GPT-4. The paper reports a split of 8,000 training, 1,000 validation and 1,000 test queries, spanning 25 domains such as arts, health and games. About 80% of queries were informational, with 10% each transactional and navigational. Each query was tagged by difficulty, genre, topic, intent and answer type so the authors could compare results by category.

That composition matters. A benchmark dominated by informational questions says more about "explain this" answers than about "which vendor should I buy" answers, which is where most B2B marketers want visibility.

How did the paper measure visibility?

The paper used two families of metrics, both relative to the other sources in the same answer:

  • Position-Adjusted Word Count measures how many words in the generated answer are attributed to a given source, weighted so that citations appearing earlier count more. A source quoted at length near the top of an answer scores highest.
  • Subjective Impression uses an LLM to judge a source's presence on seven dimensions: relevance of the cited material to the query, how much the answer relies on it, uniqueness of the material, perceived prominence of its position, perceived amount of content, likelihood a user would click it, and diversity of material presented.

Both are share-of-answer measures. They do not measure clicks, traffic, conversions or brand sentiment.

Which GEO methods did the researchers test?

The researchers tested nine rewriting methods, each applied by an LLM to the text of a source page. The table below shows each method with the Position-Adjusted Word Count reported in the paper's main results, against an unoptimized baseline of 19.3. The percentage change is calculated from those figures.

MethodWhat it changesPosition-Adjusted Word CountChange vs baseline
No optimizationNothing19.3None
Quotation AdditionAdds quotations from credible sources27.2about +41%
Statistics AdditionReplaces qualitative claims with quantitative statistics25.2about +31%
Fluency OptimizationImproves the fluency of the text24.7about +28%
Cite SourcesAdds relevant citations from credible sources24.6about +27%
Unique WordsAdds unique terms where possible20.5about +6%
Keyword StuffingAdds more keywords from the query17.7about -8%

The paper also tested Authoritative (more persuasive, confident tone), Easy-to-Understand (simpler language) and Technical Terms (more jargon). On Subjective Impression, the best method improved on the baseline by about 28%. The authors summarize the headline result as the best methods improving visibility by 41% and 28% on the two metrics respectively, which is where the widely quoted "up to 40%" figure comes from.

What does the "40%" figure actually mean?

The 40% figure is a relative improvement in one share-of-answer metric, measured in a controlled setup with five competing sources per query. It does not mean a 40% increase in AI traffic, in citations across ChatGPT, or in how often a brand gets recommended. Treat it as evidence that evidence-rich writing earns a bigger slice of a generated answer, not as a forecast for your site.

Did keyword stuffing work for AI answers?

No. Keyword stuffing was one of the weakest methods in the study. The authors wrote that while such methods are widely used for search engine optimization, they "offer little to no improvement on generative engine's responses." In the main experiment it scored below the unoptimized baseline, and on Perplexity it performed roughly 10% worse than the baseline on Position-Adjusted Word Count.

The practical reading is simple. A language model composing an answer is looking for material it can use to answer the question. Repeating the query adds no new material, so it gives the model nothing extra to quote.

Did the results hold up on a real AI search engine?

Partly, yes. The authors reran selected methods on Perplexity.ai, a live commercial answer engine at the time. On Perplexity, Quotation Addition improved Position-Adjusted Word Count by about 22%, and Statistics Addition improved Subjective Impression by about 37%, according to the paper's Perplexity results. Keyword stuffing again underperformed.

The direction of the findings survived the move from a simulated engine to a real one, while the size of the effects shifted. That pattern is a useful reminder that numbers from one engine should not be copied onto another.

Who benefited most from GEO in the study?

Lower-ranked sources benefited most. When the researchers broke results down by where a source sat in the original Google results, pages in position 5 gained dramatically while pages in position 1 often lost share. The paper reports these changes in visibility:

MethodSource ranked 1stSource ranked 5th
Cite Sources-30.3%+115.1%
Quotation Addition-22.9%+99.7%
Statistics Addition-20.6%+97.9%

The authors conclude that lower-ranked websites, which normally struggle for visibility, benefit significantly more from GEO. This is the most commercially interesting finding in the paper for smaller brands. In a generative answer, the model is choosing which material to use, not simply reading the list top to bottom, so better evidence can partly offset a weaker ranking.

There is an important caveat. The setup was zero sum among five sources. When the fifth source gained, others lost, and the first-ranked page had the most share to give up. Real AI search systems retrieve varying numbers of sources and use signals the study did not model.

Did the best tactics differ by topic?

Yes. The paper's domain analysis found that the best method depended on the type of query:

  • Authoritative tone performed well in debate, history and science queries.
  • Fluency Optimization helped most in business, science and health.
  • Cite Sources worked best for statements, facts, and law and government queries.
  • Quotation Addition was strongest for people and society, explanations and history.
  • Statistics Addition led in law and government, debate and opinion queries.

The authors also tested combinations of methods and found that pairing Fluency Optimization with Statistics Addition performed best among the combinations tried. The broad lesson is that there is no single GEO trick. The right edit depends on what kind of answer the user is asking for.

What the GEO paper did not show

The GEO paper did not show how to rank in ChatGPT, how to get recommended by Gemini, or how GEO affects traffic. Its scope is narrower than many summaries suggest. The main gaps:

It measured share of an answer, not business outcomes. There is no data on clicks, leads or revenue. A Pew Research Center analysis of 68,879 searches found users clicked a link inside a Google AI summary on just 1% of visits, so being inside an answer and receiving a visit are different outcomes.

It used a small, fixed source pool. Five sources per query, drawn from Google's top results, makes relative gains look larger than they would in a system choosing among many more pages.

The edits were machine-written. An LLM applied each rewrite. Human editing, fact-checking and original data were not part of the test.

It predates today's AI search products. The core engine was GPT-3.5-turbo. The authors themselves note that methods may need to adapt as engines evolve and that longer context windows could reduce the influence of search ranking.

It focused on on-page text. The study changed only the content of pages. It did not test third-party mentions, reviews or brand reputation, which later research suggests carry heavy weight. A 2025 study by Chen, Wang, Chen and Koudas found that AI search engines showed "a systematic and overwhelming bias towards Earned media" over brand-owned and social content, and that engines differed from each other in domain diversity, freshness and sensitivity to phrasing.

It sampled only five answers per query. AI answers vary widely from run to run. Research from SparkToro and Gumshoe across 2,961 runs found less than a 1 in 100 chance that an AI tool would return the same brand list twice. Any GEO result, including your own tests, needs repeated sampling.

How to apply the GEO paper's findings in 2026

The GEO paper is best used as a set of editorial principles, not a checklist that guarantees citations. The recommendations below are a framework based on the paper plus later public research.

1. Add evidence where it answers the question

Replace vague claims with specific, sourced statements. A hypothetical example: "our onboarding is fast" becomes "the median customer completes setup in two days, based on our Q2 onboarding data," with a link to the methodology. Only publish numbers you can back up. The paper rewarded statistics because they add usable material, and invented statistics add risk, not value.

2. Quote credible people and primary sources

Quotations performed best in the study. For a B2B page, that might mean quoting a published standard, a regulator's guidance or a named expert on your team with real credentials. Link the source so both readers and retrieval systems can verify it.

3. Cite, then write clearly

Cite Sources and Fluency Optimization both helped. Clear sentences with a direct answer at the top of each section make passages easier to extract. The Bob Builds AI guide to writing extractable content covers the structure in more detail.

4. Drop keyword repetition

Write for the question rather than the exact-match phrase. Repeating the query was the one tactic the paper showed can make things worse.

5. Match the tactic to the query type

Use the paper's domain findings as a starting hypothesis. Fact-heavy and regulatory topics may respond to citations; opinion and comparison content may respond to statistics. Then test on your own prompt set instead of assuming.

6. Work beyond your own site

Because later research points to earned media, pair on-page edits with efforts to be accurately covered by third-party sources. The comparison of AEO vs GEO vs SEO explains how on-site and off-site work fit into one program.

7. Measure with repeated sampling

Copy the paper's discipline, not just its tactics. Define a fixed prompt set, run it repeatedly across the engines your buyers use, and track the share of answers where you appear and are cited before and after a change. The guide to A/B testing for AEO walks through how to structure those comparisons.

Common mistakes when citing the GEO paper

Quoting "40% more visibility" as a promise. It is a relative gain on one metric in one setup. Presenting it as expected results for a client misrepresents the study.

Treating the nine methods as the complete GEO playbook. The paper tested on-page rewrites only. Crawler access, entity consistency and third-party coverage were outside its scope.

Stuffing pages with statistics. The benefit came from relevant, credible data. Unrelated numbers dropped into every paragraph make content worse for people and add nothing a model can use.

Assuming findings transfer unchanged to every engine. Even within the paper, effect sizes changed between the simulated engine and Perplexity. ChatGPT, Gemini, Claude and Google AI Overviews each retrieve and cite differently.

Ignoring the ranking caveat. Lower-ranked pages gained most in a five-source pool. That does not mean ranking no longer matters. Google states in its AI features guidance that the same foundational SEO practices apply to AI Overviews and AI Mode, which use "query fan-out" across related searches.

How Bob Builds AI helps

Bob Builds AI is an AEO and GEO platform and agency that helps teams apply research like the GEO paper to their own brand and measure the result. Visibility Monitoring tracks visibility rate, citation rate, competitor recommendation share, citation sources and sentiment across ChatGPT, Gemini, Claude, Perplexity, Copilot and Google AI Overviews and AI Mode, and its documentation says it measures the real chat and search interfaces rather than raw model APIs. Prompt Research surfaces the questions buyers actually ask AI assistants, which gives you the prompt set to test against. Brand Memory keeps your proof points and sourced facts in one place, so the evidence the GEO paper rewards is consistent across your content.


FAQ

What is the original GEO paper?

The original GEO paper is "GEO: Generative Engine Optimization" by Pranjal Aggarwal and co-authors from Princeton University and IIT Delhi, first posted to arXiv in November 2023 and published at KDD 2024. It coined the term generative engine optimization, introduced a 10,000-query benchmark called GEO-bench, and tested nine ways of rewriting web content to increase its visibility in AI-generated answers.

Which GEO tactic worked best in the study?

Quotation Addition worked best overall. Adding quotations from credible sources raised Position-Adjusted Word Count from a baseline of 19.3 to 27.2, roughly a 41% relative gain. Statistics Addition, Fluency Optimization and Cite Sources followed closely. On Perplexity.ai, Quotation Addition and Statistics Addition also performed strongly, although the size of the gains differed from the simulated engine.

Does keyword stuffing help with AI search visibility?

The GEO paper found keyword stuffing did not help. It scored below the unoptimized baseline in the main experiment and performed about 10% worse than baseline on Perplexity. The authors noted that methods widely used for traditional SEO offered little to no improvement in generative engine responses, because repeating query terms gives a model no new material to use.

Does the GEO paper prove a 40% increase in AI traffic?

No. The 40% figure is a relative improvement in a share-of-answer metric, measured with five competing sources per query in a controlled setup. The paper did not measure clicks, traffic, leads or brand recommendations. It shows that evidence-rich content earns a larger portion of a generated answer, which is useful directional guidance, not a traffic forecast.

Why did lower-ranked websites benefit more from GEO?

In the study, sources ranked fifth in Google gained far more visibility than first-ranked sources, with Cite Sources producing a 115.1% gain for fifth-ranked pages and a 30.3% loss for first-ranked ones. Because the answer drew on a fixed pool of five sources, better evidence let lower-ranked pages take share from higher-ranked pages. Real AI search systems use more sources and signals.

Which AI model did the GEO study use?

The main experiments used GPT-3.5-turbo as the generative engine, fed with the top five Google search results for each query and sampled five times per query at a temperature of 0.7. The authors also validated selected methods on Perplexity.ai. Because today's AI search products use newer models and different retrieval systems, the effect sizes should not be assumed to carry over directly.

Is the GEO paper still relevant in 2026?

Yes, as a foundation. Its core lessons, that specific evidence, credible quotations and clear writing help content get used in AI answers while keyword repetition does not, remain consistent with later research. Its limits also matter: it tested only on-page text, and later studies suggest earned media and third-party coverage heavily influence which sources AI search engines cite.

How should marketers test GEO tactics on their own content?

Build a fixed set of prompts your buyers realistically ask, run them repeatedly across the AI engines they use, and record how often you appear and are cited. Change one variable, such as adding sourced statistics to a page, then rerun the same prompts over several weeks. Repeated sampling matters because AI answers vary significantly from run to run.


Conclusion

The original GEO paper showed that content earns a bigger share of AI-generated answers when it carries credible quotations, relevant statistics, citations and clear writing, and that keyword stuffing does not help. It also showed that these edits helped lower-ranked pages the most, which gave smaller brands a reason for optimism.

The practical implication is to treat the paper as editorial guidance rather than a performance guarantee. Its numbers came from a five-source simulation built on GPT-3.5-turbo and measured share of an answer, not traffic or revenue. Later research adds a second layer the paper did not test: third-party coverage and earned media shape what AI search engines cite.

A sensible next step is to pick five important pages, add sourced evidence and quotations where they answer real buyer questions, and measure the change across a fixed prompt set over several weeks. Bob Builds AI can help you run that measurement across AI models and see which changes actually moved your visibility.

All posts
Findings of the original Generative Engine Optimization (GEO) research paperGEO-bench benchmarkPosition-Adjusted Word Count and Subjective Impression metricsQuotation Addition, Statistics Addition and Cite SourcesKeyword stuffing in generative engines

Don't just sit with what AI says about your brand.
Fix it now with Bob Builds.

Book a demo