The GEO Paper: What the Research Actually Found
The GEO paper is the 2024 KDD study behind the "40%" citation-visibility stat — here's what it tested, found, and where that number stops applying.
· By Rogier Bruggeman, Founder of KinetixSEO
What the GEO paper actually is
The GEO paper is the 2024 KDD study behind the "40%" citation-visibility stat — here's what it tested, found, and where that number stops applying. Its formal title is "GEO: Generative Engine Optimization," written by Aggarwal et al. and published at KDD 2024, with authors from Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi. The paper's DOI is 10.1145/3637528.3671900, and its record can be found through that identifier on a DOI resolver or the ACM Digital Library. It's the paper people mean when they cite "the 40% study" without linking it, and it's worth reading directly because the number gets repeated far more often than the conditions it was measured under.
Measure actual visibility
Want to see this on your own site?
Use your own site as the evidence. Get a free SEO and AI-citation readiness baseline, then monitor what changes.
The paper's core contribution was a new measurement methodology, not a marketing claim, built to test citation visibility instead of rank position. The authors built a benchmark called GEO-bench, made of thousands of real search queries paired with generative-engine responses, then systematically tested how rewriting the underlying content changed whether and how prominently sources got cited in the generated answer. That's a meaningfully different exercise than traditional SEO research, which typically measures rank position in a list of ten blue links. Here, the unit of measurement was visibility inside a synthesized paragraph — whether a source showed up at all, and how much of the answer's weight it carried. That framing is also why the site's own explainer on what GEO is treats it as a distinct discipline from classic SEO rather than a rebrand of it.
What the study actually tested
The researchers evaluated a set of specific, nameable content interventions rather than vague "quality" improvements, applying each one to existing web content and re-running it through generative engines to see if citation visibility changed. The interventions included:
- Adding citations — attributing claims to named sources within the text
- Adding statistics — inserting quantitative data points to support claims
- Adding quotations — including quoted statements from relevant authorities
- Keyword stuffing — repeating target terms at high density, the classic legacy-SEO tactic
- Fluency and readability edits — improving sentence-level clarity without changing substance
- Authoritative tone changes — rewriting passages to sound more definitive
Each intervention was tested in isolation so its effect could be measured independently, rather than bundling several changes together and guessing which one did the work. That separation is what makes the paper's findings usable — it tells you which specific lever moved the outcome, not just that "better content" performed better.
What moved the needle — and what didn't
Adding citations, statistics, and quotations to content increased its visibility in AI-generated answers by as much as ~40%, according to Aggarwal et al. in the original GEO paper. That's the figure that circulates widely, and it's real — but it's specifically the effect of evidentiary content: text that backs its claims with named sources, numbers, and quoted authority, not text that's merely longer or more keyword-dense. The mechanism the authors propose is plausible on its face: generative engines are synthesizing an answer from multiple sources at once, and a passage that already looks like a well-supported claim is easier to lift and attribute than one that's a bare assertion. The site's guide to writing an answer block AI engines can cite walks through what that looks like in practice.
Keyword stuffing did not produce a comparable improvement, and in some tested conditions it hurt visibility rather than helping it, per the same GEO-bench results from Aggarwal et al. This is the detail most secondhand summaries drop, and it's arguably the more useful half of the finding for anyone doing this work today: the tactic that dominated a decade of legacy SEO advice does not transfer to generative-engine visibility, at least not in the setup this paper tested. Fluency and authoritative-tone edits produced smaller, more mixed effects — directionally positive in some conditions, but nowhere near the size of the ~40% citation/statistic/quotation effect the authors measured. If you're weighing where to spend editing effort, that ordering is the paper's most concrete, actionable output.
Why the 40% figure is a ceiling, not a guarantee
The ~40% figure reported by Aggarwal et al. is an upper bound observed within one specific benchmark, one set of generative engines, and one snapshot in time — not a universal conversion rate you should expect to reproduce. Three limits matter here, and each is worth sitting with before you build a pitch or a KPI around this number:
- It's benchmark-specific. GEO-bench is a large, carefully constructed dataset, but it's still one sampling of queries and content types. A different query mix — more transactional, more technical, more local — could show a different effect size entirely.
- It's engine-and-time-specific. The generative engines tested reflect how those systems synthesized answers at the time of the study. Retrieval methods, ranking signals, and synthesis prompts inside commercial AI answer products change frequently and are not disclosed in detail by their vendors, so a result measured against 2024-era behavior is not a promise about next year's.
- It's a relative lift under experimental conditions, not a real-world booking. The ~40% describes how much more visible a rewritten passage became compared to its unmodified version inside the test harness — it says nothing about how that translates to traffic, leads, or revenue for a specific site with its own baseline content, domain authority, and competitive set.
These limits don't make the paper wrong; they make it a well-scoped academic study rather than a warranted business outcome. Treat the ~40% figure from Aggarwal et al. the way you'd treat any single experiment's effect size: as evidence a mechanism exists and is worth pursuing, not as a number you can promise a client. That distinction matters most for anyone evaluating a vendor's pitch — the site's breakdown of what GEO services actually include is a useful check against anyone quoting that figure as a delivered outcome rather than a lab result.
What this means if you're acting on it
The practical translation of the paper is a rewriting discipline: ground claims in your content with named sources, concrete numbers, and direct quotations, rather than restating the same claim in more words or repeating a target keyword. That's the specific, testable practice the paper's citation/statistic/quotation findings support — not a promise that any given page will see the ~40% lift Aggarwal et al. measured, but evidence that this particular editing pattern is worth applying to a page today, regardless of which generative engine happens to be dominant this quarter. Compared with tone polishing or fluency edits, which the paper found move the needle far less, this is where editing time is best spent first.
The paper also leaves clear gaps: it doesn't measure how sentiment inside a citation affects outcomes, and it doesn't distinguish a passing mention from becoming the answer an assistant actually recommends. Both are worth understanding separately, and both are covered in this site's piece on citation sentiment in AI search and its explainer on AI selection rate as a sharper metric than raw mention counts. For a broader run-through of the practice this paper helped name, the site's GEO guide and its practical citability checklist walk through how to apply the underlying idea to real pages.
Frequently asked questions
What is the GEO paper about?
The GEO paper is a 2024 academic study, "GEO: Generative Engine Optimization" by Aggarwal et al., that tested which specific content changes make a source more likely to be cited in AI-generated answers. It introduced a benchmark called GEO-bench and measured the effect of interventions like adding citations, statistics, and quotations against a baseline of unmodified content, comparing them to legacy tactics like keyword stuffing.
Who wrote the GEO paper?
The paper was written by Aggarwal et al., with authors affiliated with Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi, and it was published at KDD 2024, a major data-mining and applied AI research conference. Its record can be found via its DOI, 10.1145/3637528.3671900.
Does the GEO paper really prove a 40% increase in visibility?
It measured up to a ~40% relative increase in citation visibility from adding citations, statistics, and quotations, but only within its own benchmark, engine set, and time period — not as a guaranteed result anyone can reproduce. The figure, reported by Aggarwal et al., is an upper bound observed under specific experimental conditions, not a conversion rate that transfers automatically to a given site, query set, or current-generation AI engine.
Does keyword stuffing work for AI search according to this paper?
No, the paper found that keyword stuffing did not produce a comparable visibility improvement, and in some tested conditions it reduced visibility rather than helping it. This is one of the study's more concrete findings: the density-based tactic long associated with legacy search engine optimization did not transfer to generative-engine citation behavior in the setup this paper tested.
Is the GEO paper still relevant if generative engines have changed since 2024?
The specific effect sizes may not reproduce exactly on current systems, but the underlying mechanism it identified — that evidentiary content earns more citation than unsupported or keyword-dense content — remains a reasonable working assumption. Because retrieval and synthesis methods inside commercial AI answer products change often and aren't fully disclosed, the paper is best treated as a source of a testable hypothesis rather than a fixed, permanent benchmark.
Sources
- Aggarwal et al., "GEO: Generative Engine Optimization," KDD 2024 (Princeton, Georgia Tech, Allen Institute for AI, IIT Delhi) ()
Adding citations, statistics, and quotations to content increased its visibility in AI-generated answers by as much as ~40%
Measure actual visibility
See how your own site scores on SEO and AI-search visibility — free report, no signup.
Use your own site as the evidence. Get a free SEO and AI-citation readiness baseline, then monitor what changes.
Related articles
- What Is GEO (Generative Engine Optimization)?GEO & AI Search
- How to Write an Answer Block AI Engines Can CiteGEO & AI Search
- GEO Guide: How Generative Engine Optimization WorksGEO & AI Search
- GEO Services Explained: What You're Actually BuyingGEO & AI Search