GEO Study: What the Research Really Reveals About AI Search

Antonio Blago
Antonio Blago
74 readers 5.0/5 (2)
AI summary
  • Only 51.5 percent of sentences in AI search answers are fully supported by their citations according to Liu et al.
  • Up to 96 percent of users encounter at least one structurally misleading citation in AI search results per CITETRACE study.
  • The famous 40 percent GEO visibility gain measured a proxy metric on simulated engines, not real traffic or citation rank.
  • AI search systematically favors earned media from third parties over owned media according to Chen et al. research.

Generated with AI, the details are in the article.

Table of contents

GEO vendors, agencies, and freelancers promise visibility jumps in ChatGPT, Perplexity & Co., often with a catchy number: "up to 40 percent more visibility". But what does the peer-reviewed, vendor-free research really say about Generative Engine Optimization (GEO), Answer Engine Optimization (AEO), and the citation behavior of AI search systems?

This meta-study reviews 20 core academic works from 2023 to 2026, complemented by an independent field study from the Tow Center for Digital Journalism, and checks every impact claim against the primary source.

The result up front: the picture is far more sober than any marketing study. AI search cites selectively, inconsistently, and not rarely in structurally misleading ways. Users overestimate how much trust a citation deserves.

Black-hat manipulation is empirically clearly feasible, while white-hat GEO is so far surprisingly weakly supported.

Methodology of this meta-study: only peer-reviewed works and academic preprints (arXiv, publisher DOIs) were included. Vendor blogs and PR studies from SEO tool providers (Similarweb, SEMrush, Ahrefs, Profound, Ranqo, and others) are excluded a priori and serve at most as contrast. Every figure is checked against the original source; where a secondary figure and the primary source diverge, the primary source prevails. One exception is an independent, vendor-free field study by the Tow Center for Digital Journalism (source 21), which is not peer-reviewed but is flagged as an independent primary survey and used only in a complementary role.

The state of research: young, methodologically uneven, but solid

GEO is a real research field, but a young one. The solid core consists of a handful of peer-reviewed works, surrounded by a fast-growing layer of preprints. The methodological anchors are clearly identifiable: Liu, Zhang, and Liang published the first systematic test of the verifiability of generative search in 2023[1]; Gao et al. delivered the first benchmark for automatic citation evaluation with ALCE[2]. On the application side, the large-scale preference dataset Search Arena[10] and the benchmark C-SEO Bench[18] form the most solid counterweights to tool marketing.

Saturation is only partial: the period from 2023 to mid-2026 is well covered, but the 2025/26 preprint wave is highly dynamic and partly still unreviewed ("work in progress"). Methodologically clean here also means not double-counting duplicates: several works exist in parallel as an arXiv version and a publisher publication, but each is one study. The academic core corpus therefore comprises 20 studies, flanked by an independent field study from the Tow Center (source 21).

The evaluation framework: objectivity, currency, relevance, reliability

So the findings don't sit side by side arbitrarily, this meta-study reads them through four quality dimensions, modeled on classic source criticism (cf. the CRAAP framework). These four axes are the lens for everything that follows:

DimensionGuiding questionResearch finding (short)Studies
ObjectivityDoes the system select in a balanced way, or cite with a skew?Sentiment, commercial, geographic, and political bias documented throughout; the source selection itself is skewed too[8][9][11][12][14]
CurrencyAre selection and answer stable over time and prompt variants?Outputs drift over time and repetitions, the weakest axis in the entire corpus[10][12][13]
RelevanceWhich signals decide that a source is judged a good fit?Models weight query relevance strongly, stylistic credibility markers weakly[3][18]
ReliabilityDo the citations actually support the statement?Only about half of the sentences are fully supported; structural miscitations are widespread[1][16][15]

In plain terms

Without the jargon: AI search does cite sources, but those sources often do not back up what the AI claims.

  • Only about half of all sentences are actually covered by the cited source. The rest sounds backed up, but isn't.
  • Many citations point to the wrong or an off-topic source. Up to 96 percent of users encounter at least one misleading citation.
  • AI even cites texts that were themselves written by an AI as supposed evidence.

Remember: a citation is not proof. It just looks trustworthy.

How visible is your brand in ChatGPT and friends?

The AI Check shows you within minutes whether and how AI assistants mention your brand.

Start AI Check

Citation quality: selective, inconsistent, structurally misleading

The reliability axis is sobering. Liu et al. showed for four generative search engines that on average only 51.5 percent of generated sentences are fully supported by their citations, and only 74.5 percent of citations actually back the assigned statement (citation precision)[1]. Gao et al. reach a similar finding independently with the ALCE benchmark: on open-ended questions (the ELI5 dataset), even the best models fail to provide full citation support in roughly half of cases, while on narrowly defined factual questions support is higher[2].

The CITETRACE benchmark is even starker: across 11,200 queries and 112,000 answers (761,495 citation pairs), 30.6 percent of citations distort their source, 27.1 percent come from domain-inadequate sources, and up to 96 percent of users encounter at least one structurally misleading citation[16].

Important for context: this is a fresh, not yet peer-reviewed preprint. On top of that comes a self-reference problem: of the successfully checked cited sources, around 16 percent carry signs of AI-generated content (between 7 and 28 percent depending on the AI system), so AI search partly cites AI texts as evidence[15].

How consequential source selection is shows in the work on "answer bubbles": across roughly 11,000 real search queries and four systems (classic Google search, Google AI Overviews, SearchGPT, and a GPT model without search), source selection is systematically skewed. Wikipedia and long texts are overrepresented, social media and negatively framed sources underrepresented. At the same time, uncertainty markers (hedging) drop by up to 60 percent in the AI answers, while the confident language remains. So the very way sources are selected shapes which perspectives users get to see at all, and the AI sounds more certain than the evidence warrants[14].

Reliability of AI citations: only 51.5 percent fully supported sentences, 30.6 percent source-distorting citations, up to 96 percent of users encounter misleading citations

The key figures at a glance

The central figures from the primary sources, readable even without the chart (scroll sideways on small screens):

Finding Value Source
Sentences fully supported by their citation 51.5% [1] Liu et al.
Citations that truly back their statement (precision) 74.5% [1] Liu et al.
Answers without full support (open-ended questions) around 50% [2] Gao et al.
Citations that distort their source 30.6% [16] CITETRACE
Citations from off-topic sources 27.1% [16] CITETRACE
Users with at least one misleading citation up to 96% [16] CITETRACE
Cited sources with signs of AI text around 16% [15] Allaham & Diakopoulos
Drop in uncertainty markers from search integration up to 60% [14] Answer Bubbles

Field evidence: AI search often answers confidently wrong

The lab findings line up with an independent field study by the Tow Center for Digital Journalism (Columbia Journalism Review). The team gave eight AI search systems 200 verbatim excerpts from real news articles (10 articles each from 20 publishers, so 1,600 queries) and asked them to identify the headline, publisher, and original URL. The result: more than 60 percent of the answers were wrong, and mostly not with an honest admission of uncertainty but confidently wrong. The error rate varies sharply by system, from 37 percent for Perplexity up to 94 percent for Grok 3; ChatGPT Search landed at around two-thirds (134 of 200) incorrect attributions[21].

Tow Center chart: eight AI search systems and their citation behavior across 200 news excerpts, more than 60 percent wrong answers, from 37 percent for Perplexity to 94 percent for Grok 3
Each cell represents one of 200 answers per system; the more red, the more often the system was wrong. Source: Columbia Journalism Review / Tow Center for Digital Journalism (2025).

Image credit and context: The chart is from the Tow Center study by Klaudia Jaźwińska and Aisvarya Chandrasekar (Columbia Journalism Review / Tow Center for Digital Journalism, 2025) and is reproduced here for information and citation purposes with attribution. All rights to the image remain with the originator. This field study is not peer-reviewed, but an independent, vendor-free primary survey; it sits outside the academic 20-study core and complements it with field evidence.

The citation paradox: more citations ≠ more truth

If citations often fail to support, shouldn't users grow skeptical? The opposite is true. The preference dataset Search Arena (24,069 conversations, 12,652 pairwise votes, 11,650 users from 136 countries) shows: users prefer answers with more citations, even when the cited content does not support the statement at all[10]. The positive association even holds for irrelevant citations. The citation acts as a trust signal, regardless of its truth content.

On the model side this fits the finding of Wan, Wallace, and Klein: language models weight the relevance of a text to the query strongly, while largely ignoring stylistic credibility markers such as scientific references or a neutral tone[3]. Query-aligned phrasing raises the "win rate" of a piece of evidence. The result is a doubly exploitable trust signal: humans trust the quantity of citations, models trust surface relevance, both can be played without making the statement any more correct.

What AI search actually cites: concentration and bias

The objectivity axis shows skews rather than representativeness. Audits document systematic distortions in what AI search draws on as a source in the first place:

  • Li and Sinnamon find, in a 7-day audit, sentiment bias plus a commercial and geographic skew in the sources[8].
  • Kuai et al. show, for Microsoft Copilot across five languages, strong discrepancies in accuracy and attribution; in the case study on the 2024 Taiwan election, almost half of the answers contain factual errors[9].
  • Yang analyzes over 366,000 citations: only 9 percent go to news, which concentrate on a few outlets, with a pronounced liberal bias[11].
  • Zhang et al. (55,936 queries) find that 37 percent of domains appear exclusively in LLM search engines, more diverse therefore, but not more credible, neutral, or safe than classic search[12].
  • Kirsten et al. show that Google, OpenAI, and Perplexity have markedly different retrieval footprints, and that outputs vary over time and repetitions[13].

So anyone asking "how do I get into the AI answer?" is optimizing for a moving, skewed target, not for a stable, neutral index.

White-hat GEO under scrutiny: the 40-percent question

Now to the best-known figure in the whole debate. GEO stands for Generative Engine Optimization, that is, optimizing to show up in the answers of AI search systems. In their pioneering study, Aggarwal and colleagues report a visibility increase of "up to 40 percent"[4]. In marketing this number is happily quoted bare, and that is exactly what misleads.

What the "40 percent" actually measures

The number is a gain on a so-called proxy metric. A proxy metric is a stand-in measure: it does not capture the real outcome directly but only estimates it indirectly (from "proxy", a substitute). In detail:

  • What was measured is the Position-Adjusted Word Count (PAWC for short): how much of the AI answer's text comes from your source, weighted by how early it appears.
  • Not on real systems, but simulated: a rebuilt mini search engine from the top-5 Google results plus the older language model GPT-3.5-turbo, neither Google nor a real AI service.
  • No traffic, no clicks, no citation rank: it is not about real visitors, not the click rate (CTR, click-through rate, the share of users who actually click), and not how often or how prominently a source is named in the answer.
  • Only part comes from the real world: a test with 200 real queries on Perplexity gave plus 22 to 37 percent. In section 9 the authors themselves write that they did not measure ranking effects.
  • What demonstrably does nothing: keyword stuffing (cramming a text full of search terms) and sprinkling in rare words. Mostly weakly placed pages benefited (up to plus 115 percent for a page in position 5).

The key point: the apparent contradiction "GEO works (Aggarwal) but C-SEO doesn't (C-SEO Bench)" is above all a measurement difference. C-SEO stands for Conversational Search Engine Optimization, optimizing for dialogue-based AI search. One study measures a word-share estimate on a simulated engine, the other a real citation rank on a broad, standardized test. So it is not a contradiction in substance, but a comparison of apples and oranges.

The counter-test: C-SEO Bench

C-SEO Bench is the methodologically cleanest test in the entire corpus (corpus = the body of studies evaluated here). Across two tasks, namely question-answering and product recommendation, and across several subject areas, the picture is:

  • Most C-SEO methods have no measurable effect, and several even worsen the citation rank markedly.
  • Clearly more effective is classic, retrieval-oriented SEO (SEO stands for Search Engine Optimization). "Retrieval-oriented" means making sure your source is found by the system at all and pulled into the answer context of the language model (LLM, short for Large Language Model)[18].
  • The more competitors apply the same tricks, the smaller the gains. The study calls the effect "congested and zero-sum", that is, overcrowded and a zero-sum game where in the end no one gains anything.

What does work reliably: earned media

What is robustly supported, by contrast, is something else: Chen et al. (University of Toronto) show that AI search systematically and very clearly favors so-called earned media. Earned media are credible pieces that third parties publish about you (such as press articles or expert posts), as opposed to your own web pages (owned media) and social-media posts. On top of that comes a "big-brand bias", a built-in preference for large, well-known brands[17]. Both findings are stated verbatim in the study's abstract. Caution is warranted only with the tone of the work (its title is "How to Dominate AI Search"); the measurements themselves are solid.

In practice: how to act on these findings

  • Treat the "40 percent" as a lab proxy, not a promise. Measure success by real citations in ChatGPT, Perplexity, and Google's AI answers, not by a word share.
  • Skip keyword stuffing and trick texts: they demonstrably don't work and can even worsen your citation rank.
  • Invest in earned media: mentions, reviews, and expert articles on authoritative third-party sites count more than your own landing page.
  • Make content citable: clear structure, unambiguous statements, sourced facts, so the system finds you and pulls you into the answer.
  • Differentiate on substance: when everyone runs the same tactic, the effect shrinks ("congested and zero-sum"). Real, verifiable substance stays the edge.
GEO impact claimed vs. demonstrated: plus 40 percent word-share proxy on a simulated engine (single study) against congested and zero-sum, mostly ineffective on a broad benchmark

How to appear in AI answers: the newer evidence

If classic C-SEO tricks barely move the needle (Puerto et al.[18]) and the famous "40 percent" from Aggarwal et al. is only a proxy on a simulated engine[4], the practical question remains: what makes a source appear in AI answers at all? Two newer works shift the focus away from the "trick" toward content and citability.

Tian et al. (AgentGEO) turn the GEO question around: instead of promising "more visibility" wholesale, they diagnose why a specific document is not cited and repair it in a targeted way. Their agentic system achieves more than a 40 percent relative increase in citation rate with only around 5 percent of the content changed, while generic standard methods manage only about 25 percent[19]. Unlike the Aggarwal figure, this is a directly measured citation rate, not a word-share proxy. Two caveats remain: it is a fresh, not yet peer-reviewed preprint (March 2026), and the authors themselves show that generic optimization can harm long-tail content and that some documents cannot be repaired by text edits alone, an open question of fair visibility.

Chen et al. (CC-GSEO-Bench) provide the matching measurement instrument from the content creator's perspective. Their benchmark couples a large dataset (over 1,000 source articles, over 5,000 query-article pairs) with a creator-centered evaluation framework and measures a source's influence across five dimensions: exposure, faithful credit (correct attribution), causal impact, readability and structure, and trustworthiness and safety[20]. The message: being cited is multi-dimensional, and whether a source truly exerts influence can only be judged once real causal contribution is separated from mere mention.

Read together with the rest of the corpus, a consistent picture of the levers emerges, in descending order of evidential strength: most strongly and most cleanly supported experimentally is a source's retrieval rank in the model context, whether and where it is presented to the model at all (Puerto[18]). Broadly documented next is the lever of earned media and brand signals, authoritative third-party sources rather than self-promotion (Chen et al. / Toronto[17]). On the content side, it pays to be deliberately citable rather than to play stylistic tricks (AgentGEO[19], CC-GSEO-Bench[20]). And all of that is worth only as much as it is correctly attributed, since nearly a third of citations distort their source (CITETRACE[16]).

Black-hat: manipulation is empirically feasible

Ironically, the manipulation literature is currently more robust than the white-hat literature. Several controlled experiments show that AI search can be deliberately influenced:

  • Pfrommer et al. demonstrate on the RAGDOLL dataset (1,147 real product web pages) that a "tree-of-attacks" technique reliably promotes low-ranked products to the top, with transfer to Perplexity[5].
  • Kumar and Lakkaraju show that a "strategic text sequence" makes a target product the top recommendation significantly more often[6].
  • Nestaas et al. demonstrate "preference manipulation attacks" against Bing, Perplexity, and GPT-4 and Claude plugins, with a prisoner's-dilemma logic that degrades the overall output for everyone[7].

The line between legitimate optimization and adversarial manipulation is technically thin. That is exactly why black-hat is not a strategy recommendation: it works, but it is fleeting due to provider patches and carries substantial reputational and legal risks.

The 20 core studies at a glance

#Study (year)TypeKey findingStatus
1Liu et al. (2023)AuditOnly 51.5% of sentences fully supported by citationsEMNLP Findings
2Gao et al. / ALCE (2023)BenchmarkBest models lack full citation support in ~50% of casesEMNLP
3Wan et al. (2024)ExperimentModels weight relevance, ignore style/toneACL
4Aggarwal et al. (2024)Experiment (proxy/sim.)+40% on word-share proxy (PAWC); ranking not measuredKDD
5Pfrommer et al. (2024)AdversarialTree-of-attacks promotes low-ranked productsEMNLP
6Kumar & Lakkaraju (2024)Experiment (small)Strategic text sequence lifts target productPreprint
7Nestaas et al. (2024)AdversarialPreference manipulation against Bing/Perplexity/pluginsPreprint
8Li & Sinnamon (2024)AuditSentiment, commercial, and geographic biasASIS&T
9Kuai et al. (2025)AuditLanguage-dependent errors and attribution (Copilot)New Media & Society
10Miroyan et al. / Search Arena (2025)Preference dataMore citations = more approval, even without supportICLR
11Yang (2025)Secondary analysis9% news, outlet concentration, liberal biasPreprint
12Zhang et al. (2025)Empirical37% of domains LLM-exclusive, but not more crediblePreprint
13Kirsten et al. (2025)ComparisonDifferent retrieval footprints, output driftPreprint
14Huang et al. / Answer Bubbles (2026)AuditSource-selection bias; hedging drops by up to 60%Preprint
15Allaham & Diakopoulos (2026)Audit~16% of cited sources AI-generatedAAAI
16Seo et al. / CITETRACE (2026)Benchmark30.6% distort the source; up to 96% of users affectedPreprint (WIP)
17Chen et al. / Toronto (2025)ExperimentEarned-media and big-brand biasPreprint
18Puerto et al. / C-SEO Bench (2025)BenchmarkC-SEO mostly ineffective/negative; position dominatesNeurIPS D&B
19Tian et al. / AgentGEO (2026)Experiment (benchmark)Targeted repair: +40% citation rate at 5% content change (baseline 25%)Preprint
20Chen et al. / CC-GSEO-Bench (2025)BenchmarkContent-centric measure of source influence (5 dimensions)Preprint

Deliberately not in this core are industry-aligned and vendor works: the commercially marketed GEO-16 framework, statistical industry frameworks with affiliation not primarily verified, and the platform study of a GEO tool provider (Ranqo). They partly overlap thematically with the big-brand bias, but as vendor or borderline sources they are kept separate and serve only as contrast.

Transparency note: my own, complementary research

For context, two studies of my own that are deliberately not part of the vendor-free 20-study core. They come from me and therefore carry a conflict of interest. They do not contradict the core findings, they illustrate them on real data:

  • LLM visibility study (144,000 data points, 48 brands, 100 prompts, three models): which brand a model names varies strongly by model. AG1, for instance, appears in 62% of Gemini answers but only 5% with GPT-4o. This matches the model-dependent retrieval drift in Kirsten et al.[13] Code and data are open (MIT license).
  • E-commerce GEO case study (PURELEI) (117 prompts across ChatGPT, Perplexity, Google AI Overviews, and Claude): shows that GEO visibility can translate measurably into revenue, here around EUR 24,473 in annual ChatGPT revenue at a 2.98% conversion rate. A real-world demonstration that earned-media and brand signals (columns G, H, I in the matrix) work commercially too.

Methodological note: both are my own, published studies, not peer-reviewed third-party research. They complement the core as a practical illustration, they do not replace it.

What brands and SEO teams should take from the evidence

Condense the 20 studies into concrete levers and a clear pattern emerges of which approaches are supported and which are overrated. The evidence matrix below shows how strongly individual GEO levers are supported across the studies:

Evidence matrix: retrieval position and adversarial injection experimentally well supported, earned media broad but correlational, stylistic tricks and llms.txt weak to negative

The same matrix as a searchable table, readable even without the chart (scroll sideways on small screens). Columns A to J stand for the individual studies, the legend below the table resolves them:

Lever \ Study ABCDEFGHIJ
Retrieval / context position + (D)++ (E)+ (E)··+ (E)··+ (D)·
Earned media / brand ······++ (C/D)+ (C)+ (D)+ (D)
Metadata & freshness ······+ (D)+ (C)+ (C)+ (D)
Structured data / schema 0 (E*)0 (E)····+ (D)+ (C)+ (C)·
Semantic HTML ······+ (D)+ (C)+ (C)·
Adding statistics + (E*)0/- (E)···0 (E)+ (D)·+ (C)·
Adding citations / quotes + (E*)0/- (E)···- (E)·+ (C)+ (C)+ (C)
Stylistic edits + to 0/- (E*)-- (E)···- (E)····
llms.txt / LLM guidance ·0 (E)········
Adversarial prompt injection ·(context)++ (E*)++ (E*)++ (E*)·····

Studies (columns): A = Aggarwal et al., GEO ([4]) · B = Puerto et al., C-SEO Bench ([18]) · C = Pfrommer et al. ([5]) · D = Kumar & Lakkaraju ([6]) · E = Nestaas et al. ([7]) · F = Wan et al. ([3]) · G = Chen et al. / Toronto ([17]) · H = GEO-16 framework (industry, not in core) · I = citation-absorption (borderline preprint, not in core) · J = audit group, observational audits of source bias ([1] [8] [11] [12] [15])

Effect: ++ strongly positive · + positive · 0 no robust effect · - negative · -- significantly negative · · not studied. Evidence type: E = experimental/significant · C = correlational · D = descriptive · E* = small/synthetic sample or manipulation.

The empty cells are a finding in themselves: outside industry-aligned, correlational sources, semantic HTML, structured data, and llms.txt are practically unstudied. Clean, experimental causality exists almost only for retrieval position and adversarial injection.

  • Mirror strong effectiveness promises against the benchmarks. Anyone promising "+X percent AI visibility" should be able to explain whether that is real citation rank or a word-share proxy, and whether it was checked against C-SEO Bench and Search Arena[18][10].
  • Retrieval rank and source properties beat micro-optimization. Classic authority and earned media remain the most solid levers, while stylistic edits are empirically weak[18][17].
  • Being cited is not being cited correctly. With 30.6 percent of citations distorting, you need monitoring, not just placement[16][1].
  • Currency is the weakest axis. Every finding is a fast-decaying snapshot, so measure repeatedly rather than once, ideally with confidence intervals[10][13].
  • Black-hat works, but it is a risk. Provider patches, reputational, and legal risks make manipulation a dead end[5][6][7].

How visible is your brand really?

Instead of guessing whether your brand is visible in AI search and in the market: the free Brand Radar shows you monthly brand search volume, 3- and 12-month trends, plus winners and losers across hundreds of D2C brands in the DACH region. Measure visibility, don't assume it.

Start the Brand Radar

Conclusion

The vendor-free balance can be put in one sentence: AI search cites selectively, inconsistently, and not rarely in structurally misleading ways; users overestimate how much trust a citation deserves; black-hat manipulation is empirically feasible; and white-hat GEO is so far weaker supported than tool marketing suggests, the much-cited "40 percent" is a proxy metric on a simulated engine, not a measured citation rank. For a serious GEO strategy this means: evidence before tool promises, robust levers (retrieval rank, earned media, machine-readable structure) before stylistic tricks, and repeated measurement rather than a single snapshot.

Want to make this concrete for your brand? Start with the free Brand Radar, or book a free initial consultation directly, where we assess your GEO and SEO visibility on the evidence.

AB

Author

Antonio Blago – Neuro-SEO System®

SEO freelancer and creator of the Neuro-SEO System®. Antonio helps brands with visibility in classic and AI-driven search and reviews GEO research consistently vendor-free, evidence before tool promises.

📧 info@antonioblago.com 📞 +49 163 698 0955 📍 Hohenzollernstraße 143, 56058 Koblenz

Abbreviations and key terms explained

So this piece stays readable without prior knowledge, here are the most important terms in one sentence each:

  • GEO (Generative Engine Optimization): optimizing content so it is named and cited in the answers of generative AI systems (ChatGPT, Perplexity, and the like).
  • AEO (Answer Engine Optimization): almost identical to GEO, focused on answer engines that deliver a direct answer instead of a list of links.
  • C-SEO (Conversational SEO): SEO for dialog-based AI search; in this corpus the umbrella term for the tested content tricks.
  • GSE (Generative Search Engine): an AI search system that retrieves sources and formulates a generated answer from them.
  • LLM (Large Language Model): the large language model, the AI behind the answer generation.
  • RAG (Retrieval-Augmented Generation): a method in which the model retrieves matching sources before answering (retrieval) and folds them into the answer.
  • Retrieval / retrieval rank: the fetching of sources and the position at which a source is handed to the model as context; per the evidence the strongest visibility lever.
  • Earned media: reach through authoritative third-party sources (press, specialist sites, Wikipedia) rather than owned channels or paid ads.
  • PAWC (Position-Adjusted Word Count): the proxy metric from the Aggarwal study, a position-weighted word share of a source in the AI answer, not measured traffic and not a citation rank.
  • Proxy metric: a stand-in measure that only approximates the actual goal (for example traffic); strong promises built on a proxy should be read with caution.
  • STS (Strategic Text Sequence): an inserted text sequence that nudges a model to prefer a particular product, a manipulation technique.
  • Hedging: linguistic uncertainty markers (possibly, according to the source); these drop in AI answers, which then sound more certain than the evidence warrants.
  • CTR (Click-Through Rate): the share of users who click on a displayed result. For empirical per-position values, see the Keyword Study 2026.
  • Benchmark: a standardized test dataset against which methods are measured objectively and comparably (for example ALCE, C-SEO Bench, CC-GSEO-Bench).
  • Preprint: a study published in advance (usually on arXiv), not yet through full peer review.
  • Peer review: evaluation by independent expert peers before publication, a quality filter that preprints have not yet passed.
  • D2C (Direct-to-Consumer): brands that sell directly to end customers without intermediaries.

Sources

  1. Liu, N. F., Zhang, T. & Liang, P. (2023) – Evaluating Verifiability in Generative Search Engines. Findings of EMNLP 2023. arxiv.org/abs/2304.09848
  2. Gao, T., Yen, H., Yu, J. & Chen, D. (2023) – Enabling Large Language Models to Generate Text with Citations (ALCE). EMNLP 2023. arxiv.org/abs/2305.14627
  3. Wan, A., Wallace, E. & Klein, D. (2024) – What Evidence Do Language Models Find Convincing? ACL 2024. arxiv.org/abs/2402.11782
  4. Aggarwal, P. et al. (2024) – GEO: Generative Engine Optimization. KDD 2024, DOI 10.1145/3637528.3671900. arxiv.org/abs/2311.09735
  5. Pfrommer, S., Bai, Y., Gautam, T. & Sojoudi, S. (2024) – Ranking Manipulation for Conversational Search Engines. EMNLP 2024. arxiv.org/abs/2406.03589
  6. Kumar, A. & Lakkaraju, H. (2024) – Manipulating Large Language Models to Increase Product Visibility. arxiv.org/abs/2404.07981
  7. Nestaas, F., Debenedetti, E. & Tramèr, F. (2024) – Adversarial Search Engine Optimization for Large Language Models. arxiv.org/abs/2406.18382
  8. Li, A. & Sinnamon, L. (2024) – Generative AI Search Engines as Arbiters of Public Knowledge: An Audit of Bias and Authority. Proc. ASIS&T 2024. DOI 10.1002/pra2.1021
  9. Kuai, J., Brantner, C., Karlsson, M., Van Couvering, E. & Romano, S. (2025) – AI chatbot accountability in the age of algorithmic gatekeeping. New Media & Society. DOI 10.1177/14614448251321162
  10. Miroyan, M., Wu, T.-H. et al. (2025) – Search Arena: Analyzing Search-Augmented LLMs. ICLR 2026. arxiv.org/abs/2506.05334
  11. Yang, K.-C. (2025) – News Source Citing Patterns in AI Search Systems. arxiv.org/abs/2507.05301
  12. Zhang, P., Ye, Q., Peng, Z., Garimella, K. & Tyson, G. (2025) – Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines. arxiv.org/abs/2512.09483
  13. Kirsten, E. et al. (2025) – Characterizing Web Search in the Age of Generative AI. arxiv.org/abs/2510.11560
  14. Huang, M., Goyal, A., Saha, K. & Chandrasekharan, E. (2026) – Answer Bubbles: Information Exposure in AI-Mediated Search. arxiv.org/abs/2603.16138
  15. Allaham, M. & Diakopoulos, N. (2026) – Synthetic Sources? Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources. AAAI 2026. arxiv.org/abs/2605.23684
  16. Seo, Y., Jeong, W., Kim, E., Jang, H. & Lee, D. (2026) – Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs (Work in Progress). arxiv.org/abs/2605.28565
  17. Chen, M., Wang, X., Chen, K. & Koudas, N. (2025) – Generative Engine Optimization: How to Dominate AI Search. arxiv.org/abs/2509.08919
  18. Puerto, H., Gubri, M., Green, T., Oh, S. J. & Yun, S. (2025) – C-SEO Bench: Does Conversational SEO Work? NeurIPS 2025 Datasets & Benchmarks. arxiv.org/abs/2506.11097
  19. Tian, Z., Chen, Y., Tang, Y., Liu, J. & Jia, R. (2026) – Diagnosing and Repairing Citation Failures in Generative Engine Optimization (AgentGEO). arxiv.org/abs/2603.09296
  20. Chen, Q., Chen, J., Huang, H., Shao, Q., Chen, J., Hua, R., Xu, H., Wu, R., Chuan, R. & Wu, J. (2025) – CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines. arxiv.org/abs/2509.05607
  21. Jaźwińska, K. & Chandrasekar, A. (2025) – We compared eight AI search engines. They’re all bad at citing news. Columbia Journalism Review / Tow Center for Digital Journalism. cjr.org

Free SEO tools

Keyword ranking, SSR text check, sitemap audit and more, right in the browser and without signup.

See SEO tools