What the Princeton GEO Study Actually Found (and How to Apply It)
The real methodology and numbers behind the most-cited GEO research, what actually moved citation rate, and what backfired.
The 2024 Princeton "GEO: Generative Engine Optimization" study tested nine content-optimization methods across roughly 10,000 queries. Adding statistics and adding quotations from authoritative sources were the strongest single levers, keyword stuffing measurably hurt visibility, and lower-ranked pages benefited disproportionately more than already-top-ranked ones.
What was actually tested
The paper, by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, was published at KDD 2024, one of the more respected venues for this kind of applied research. The team built a benchmark called GEO-bench, ran it across roughly 10,000 diverse queries, and tested nine distinct content-optimization strategies against two metrics: Position-Adjusted Word Count (how much of the generated answer's text traces back to your content, weighted by where it appears) and Subjective Impression (a separate rating of how favorably the response comes across).
The nine methods tested were citations, statistics addition, quotation addition, technical terminology, fluency optimization, keyword stuffing, authority claims, unique content, and simplification.
The two strongest single levers
Statistics addition was the single strongest method tested: +41% on Position-Adjusted Word Count and +37% on Subjective Impression. Adding quotations from authoritative sources was close behind. Both work for the same underlying reason our own named-sources check is built around: a claim attached to a specific number or a specific attributed source is easier for a model to lift out of context and repeat confidently. A vague claim like "experts agree" gives a model nothing concrete to quote.
Combining the two top strategies (fluency optimization plus statistics addition) outperformed any single method by more than 5.5%, meaning the gains from good individual tactics do stack, at least to a point.
What backfired
Keyword stuffing, the oldest trick in classical SEO, actively decreased visibility in generative engine responses. This is a meaningful signal about how differently these two disciplines actually work at the technical level: a tactic optimized for a ranking algorithm reading dozens of keyword-density signals can read as noise, or even as a negative trust signal, to a model that's trying to extract a coherent answer rather than match a query string.
The equity finding: smaller sites benefit more
One of the more surprising results: a fifth-ranked page in the study saw a 115.1% increase in AI visibility from adding cited sources alone, while the already-top-ranked page in the same scenario lost 30.3% of its share of the generated response. In plain terms, if you're not already the most authoritative source on a topic, source attribution and statistics are where you have the most room to move the needle, an already-dominant page has less to gain and can even lose relative share as competitors close the gap.
| Finding | Direction | Note |
|---|---|---|
| Statistics addition | Strong positive (+41% PAWC, +37% Subjective Impression) | Strongest single method tested |
| Quotation addition | Strong positive | Close second to statistics |
| Fluency + Statistics combined | Positive, +5.5% over best single method | Gains stack |
| Keyword stuffing | Negative | Actively hurts visibility |
| Lower-ranked page + citations | Strong positive (+115.1% in one scenario) | Smaller sites gain disproportionately |
| Top-ranked page, same scenario | Negative (-30.3% share) | Already-dominant pages have less to gain |
How to apply this without guessing
This maps directly to two checks already in our audit: named, linked sources instead of vague appeals ("studies show"), and concrete numbers instead of qualitative claims ("grew significantly" versus "grew 34%"). Neither requires new content strategy, most pages already have the real numbers and real sources, they're just written as adjectives instead of attributed facts. Run the full AEO & GEO audit and both show up as scored, specific checks rather than vague advice, so you know exactly which sentence to rewrite.
FAQ
Do I need an academic-grade statistic, or does a real number from my own data count?
Your own real, verifiable number counts. The study rewards concrete, attributable figures over vague qualitative claims, not any particular source of the number.
Does this mean I should cram every paragraph with quotes and numbers?
No, the study measured statistics and quotations added where they genuinely support a claim, not volume for its own sake, and keyword-stuffing-style overuse of any tactic measurably hurt results.
Is GEO actually different from AEO based on this paper?
Not meaningfully at the technical level, the methods that worked (attribution, concrete numbers, clear writing) are the same ones AEO already optimizes for. See our AEO vs SEO vs GEO post for the full breakdown.