All articles
Fundamentals

GEO Evidence Audit 2026: What 45 Studies Actually Prove

A July 2026 critical survey reviewed 45 GEO studies and found no technique shows a stable cross-platform causal effect on organic discoverability. We break down what the research actually proves: the +41% quotation lift is position-adjusted word count (19.3 to 27.2), not clicks or retrieval; Kumar et al. (arXiv:2606.20065) benchmarked 100K+ prompts showing a 3-tier brand ladder (73% / 44% / 11%); and indexing, not formatting, is the real bottleneck. With FAQ schema and the Princeton KDD 2024 baseline.

11 min read·Updated 2026-08-19

Generative Engine Optimization has a measurement problem, not a hack problem. Every week a new agency claims it "got a client cited in ChatGPT" through some formatting trick. A July 2026 critical survey reviewed 45 GEO studies and found the tactical playbook mostly does not hold up under scrutiny. This guide separates what the research actually proves from what it does not — so you can invest in the levers with reproducible evidence, not the ones that sell workshops.

The two studies that matter most for 2026 practitioners are the original Princeton GEO benchmark (Aggarwal et al., KDD 2024) and the new large-scale benchmark from Kumar et al. (arXiv:2606.20065, June 2026). Layered on top of both is the critical survey of 45 studies (Martinez, arXiv:2607.14035, July 2026) that reframes the whole field. Read together, they tell a clearer story than any single vendor case study.

What the 2026 evidence says: The Princeton quotation tactic lifted position-adjusted word count from 19.3 to 27.2 (+41%) — a share-of-answer-text gain, not clicks. The July 2026 survey of 45 studies found no technique with a stable cross-platform causal effect on organic discoverability. Kumar et al. benchmarked 100,000+ prompt responses across 100+ brands and found a 3-tier brand ladder: household names cited in 73% of answers, mid-market 44%, niche 11%. 78% of citations go to corporate websites; ranked "best-of" listicles are the top format at 21%.

Tracing the famous "+40%" to its source

Almost every GEO pitch deck quotes "up to 40% more visibility." The July 2026 critical survey traces that number to its origin and shows exactly what it measures. In Aggarwal et al. (KDD 2024), the outcome variable is position-adjusted word count — the share of the generated answer attributed to your source, weighted by where it appears in the answer.

MetricBeforeAfterWhat it means
Quotation tactic (Princeton)19.327.2+41% relative share of answer text
ScopeFixed testbedFixed testbedSource already in model context
Does NOT measureClicks, retrieval, or discoverability

The critical survey is explicit: this "+41%" does not mean 40% more readers will click, nor that a page gains 40% in retrieval probability. It means that, in this testbed, a source already provided to the generator receives a larger position-weighted share of attributed text. The generalized claim that "GEO increases visibility by 40%" is listed among the claims the review rejects as unsupported. The effect is real; the interpretation sold around it is not.

The 2026 critical survey: 45 studies, one reframe

Martinez (arXiv:2607.14035, July 2026) published the first critical survey of generative engine optimization, reviewing 45 studies from November 2023 to July 2026. Its central reframing: visibility in generative engines is not a single ranking task but a stochastic, partially observable pipeline. It separates seven outcome variables rather than one score — activation, crawling and indexing, retrieval, reranking and context allocation, generation and citation, absorption and fidelity, and finally attention, clicks and conversion.

"Claims about GEO return on investment clearly outstrip the academic evidence. A tactic that improves one stage can be irrelevant, or harmful, at another. This is why a single 'AI visibility percentage' is a misleading KPI: it collapses seven different failure points into one number that cannot tell you which one is broken."
— Synthesis of Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of GEO" (arXiv:2607.14035, July 2026)

Two findings matter most for budget allocation. First, topical relevance and where content sits within the retrieved context are the only levers with reproducible evidence behind them. The generic formatting heuristics that fill most GEO checklists — bullet structuring, adding statistics, writing in a Q&A format — transfer poorly across platforms and frequently fail to replicate. Second, day-to-day source overlap is low: the same query returns a Jaccard similarity of just 0.34–0.42 across repeated runs, so a single before/after screenshot proves nothing without repeated sampling.

Large-scale benchmark: the 3-tier brand ladder (Kumar et al., 2026)

While the critical survey cools expectations on discoverability, Kumar et al. (arXiv:2606.20065, June 2026) provide the first large-scale empirical baseline for measuring GEO. They analyzed 100,000+ prompt responses across 100+ brands tracked between March and May 2026. The results quantify a clear brand-stature ladder and the formats that win citations:

FindingMetricImplication
Brand-stature ladder73% / 44% / 11%Household names, mid-market, and niche brands appear in that share of relevant answers — about 30 points per step
Source of citations78% corporate sitesAmong non-corporate sources, YouTube leads ahead of Reddit, editorial media, and Wikipedia
Top content format21% of citationsRanked "best-of" listicles are the single most-cited page type
Sentiment instability6.7× flip rateWhether a brand is framed positively or negatively flips 6.7× more often than whether it is mentioned at all

Source: Kumar et al., "Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines," arXiv:2606.20065 (June 2026).

What this means for your GEO program

The evidence does not say "stop doing GEO." It says scope it to the stage you can actually control. Three takeaways:

  1. 1.
    Treat indexing as the real bottleneck. The July 2026 survey is blunt: no technique yet shows a stable effect on organic discoverability. Get into every engine index first — see the get-indexed-by-AI-search-engines guide. Optimization only pays off once you are in the retrieval pool.
  2. 2.
    Optimize for citation, measured correctly. The Princeton lifts are real for already-retrieved pages: expert quotations +41%, statistics +33%, fluency +29%, citations +28%. But measure with repeated sampling across prompts, not single screenshots — day-to-day overlap is only 0.34–0.42.
  3. 3.
    Build entity authority, not just on-page tricks. Kumar et al. show a 3-tier brand ladder and that 78% of citations go to corporate sites while best-of listicles win 21% of citations. Topical relevance and early context placement are the only levers with reproducible evidence — earn them through real coverage, not keyword stuffing (which costs −8%).

Frequently asked questions

What did the 2026 GEO critical survey actually find?

The July 2026 survey (Martinez, arXiv:2607.14035) reviewed 45 GEO studies and concluded that GEO techniques reliably change how an already-retrieved page is cited, but no reviewed technique shows a stable, cross-platform causal effect on organic discoverability. It is the first systematic audit of the field and a useful corrective to overclaimed tactical wins.

Is the Princeton "+41%" number fake?

Not fake, but frequently misread. It measures position-adjusted word count — the share of the AI answer attributed to a source, weighted by position — rising from 19.3 to 27.2. That is a real, reproducible citation-gain signal for already-retrieved pages. It is not a click gain or a retrieval gain, which is the part most pitches get wrong.

What is the best content format for AI citations?

Kumar et al. (2026) found ranked "best-of" listicles are the most-cited format at 21% of all citations, ahead of other structures. Combined with the Princeton finding that statistics (+33%) and expert quotations (+41%) lift citation, a well-sourced listicle with named data is a strong GEO template.

Should small or niche brands bother with GEO?

Yes, but with realistic expectations. The 2026 large-scale benchmark shows niche brands appear in only 11% of relevant answers versus 73% for household names — about 30 points per tier. The lever with reproducible evidence is topical relevance and early context placement, which smaller brands can win on specific subtopics even without broad authority.

References: Aggarwal, P., Dugan, L., et al. "GEO: Generative Engine Optimization." arXiv:2311.09735, KDD 2024. · Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of GEO." arXiv:2607.14035 (July 2026) — review of 45 studies, Nov 2023–Jul 2026. · Kumar et al. "Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines." arXiv:2606.20065 (June 2026) — 100K+ prompt responses across 100+ brands. · Princeton GEO benchmark (GEO-bench, 10,000 queries × 9 datasets). · Capston AI, "The GEO Evidence Audit: What 45 Studies Actually Prove" (2026 synthesis of the critical survey). · AI Eating the World, "GEO Has a Measurement Problem, Not a Hack Problem" (2026).

Want to check your site's GEO readiness?

Run the 27-point GEO audit