How LLMs Rank Brands: A Statistical Study on AI Visibility with Claude, GPT-4o and Gemini

Antonio Blago
Antonio Blago
0 readers 5.0/5 (1)

How do AI models like ChatGPT, Claude and Gemini decide which brands to recommend? We developed an open-source framework to measure exactly that, with 48 supplement brands, 3 LLMs and 144,000 data points.

The results reveal massive differences between the models and provide concrete insights for brands looking to improve their AI visibility.

The Problem: AI Search is a Black Box

When a user asks ChatGPT "Which supplement brands are good in Germany?", the model delivers a list of recommendations. But unlike traditional SEO, where you can see rankings in Google Search Console, LLMs offer no transparency about how brands are ranked.

We developed the LLM Brand Visibility and Ranking Framework, a reproducible, statistical methodology for measuring brand visibility across multiple AI models.

Study Design

Our methodology follows scientific standards with statistical robustness:

  • 3 LLM models: Claude (Anthropic), GPT-4o (OpenAI), Gemini 2.5 Flash (Google)
  • 48 supplement brands from the Visibly AI Brand Radar (German market)
  • 100 generic prompts in 2 clusters (Informational, Commercial), no brand names in the prompts
  • 10 repetitions per prompt per model (Temperature=0.7) for variance measurement
  • 144,000 evaluations in total (100 prompts x 3 models x 10 runs x 48 brands)
  • Buyer persona: 25 years old, goes to the gym, nutrition-conscious, asks from a first-person perspective

All prompts explicitly ask for brand recommendations ("Which brands…"), but never mention a specific brand. Brands are only identified in the output, using 3-layer detection: Exact Match, Domain Match and Fuzzy Match.

Result 1: ESN and AG1 Dominate AI Recommendations

Brand Visibility Ranking

The top brands by mention rate across all three models:

Rank Brand Mention Rate Top-3 Rate
1 ESN 28,3% 9,1%
2 AG1 26,0% 3,4%
3 Sunday Natural 13,6% 2,2%
4 Foodspring 13,1% 3,2%
5 Ritual 4,6% 0,4%

ESN (Elite Sports Nutrients) leads with a mention rate of 28,3%, meaning the brand appears in roughly every 3.5th AI response about supplements. AG1 (Athletic Greens) follows with 26%, but with a dramatically different distribution across models.

Result 2: The Models Disagree Massively

Modellvergleich

The three models show fundamentally different behavior:

  • Gemini is the most brand-heavy (4.7% average mention rate), recommending almost 5x more brands than GPT
  • Claude falls in the middle (2.2%), balanced between brand mentions and generic advice
  • GPT-4o is the most conservative (1.0%), preferring generic supplement advice over specific brand recommendations

Result 3: AG1 has a 57% spread between models

Brand Volatilität

The most striking finding is brand volatility – how differently each model treats the same brand:

Brand Claude Gemini GPT Spread
AG1 11,2% 62,0% 4,9% 57,1%
Sunday Natural 16,6% 21,7% 2,5% 19,2%
Foodspring 20,5% 14,6% 4,3% 16,2%
ESN 28,0% 36,1% 20,8% 15,3%

AG1 appears in 62% of all Gemini responses, but in only 5% of GPT responses. That is a 12-fold difference for the same brand using the same prompts. For brands, this means that GEO (Generative Engine Optimization) cannot be a one-size-fits-all solution. Each model requires its own strategy.

Surprise: More Nutrition only in 6th place

A particularly interesting result: More Nutrition lands in only 6th place with a 4.5% mention rate, even though the brand is one of the top-selling in the German supplement market. In our More Nutrition revenue and SEO analysis, we showed that the brand achieves over 823,000 monthly branded searches and an estimated 800 million EUR in annual revenue.

So why does More Nutrition rank so low in AI visibility? One possible explanation is the target audience: Our buyer persona is a 25-year-old gym-goer who asks generically about supplement brands. More Nutrition positions itself heavily through influencer marketing (particularly via founder Christian Wolf), which may be less prominent in LLMs' training data than the traditional SEO presence of ESN or AG1.

This shows: High search volume and brand awareness do not automatically guarantee high AI visibility. GEO requires different signals than traditional SEO or social media marketing.

Methodology: How We Measured This

Our framework uses a robust statistical approach:

  1. Generic prompts only, no brand names in the prompts. Brands are only identified in the LLM output.
  2. Buyer persona, all prompts simulate a real user: 25 years old, gym-goer, health-conscious, first-person perspective.
  3. 3-layer brand detection, Exact Match (Regex with word boundaries), Domain Match (brand URL in the response), Fuzzy Match (thefuzz library for typos).
  4. Statistical tests, Fisher Exact Test for mention rates, Mann-Whitney U for rankings, Benjamini-Hochberg FDR correction.
  5. Bootstrap confidence intervals, 5.000 resamples for all metrics.
  6. Power analysis, Monte-Carlo simulation confirms >90% power at 10 runs x 100 prompts.

Power Analysis: How Many Runs Do You Need?

A common question in LLM research: How many repetitions do you need for statistically reliable results? We ran Monte Carlo simulations to find out.

Configuration Observations Power (5pp Effect) Cost (3 Models)
10 Runs x 50 Prompts 1.500 34% (too low) ~9 EUR
30 Runs x 50 Prompts 4.500 80% (minimum) ~28 EUR
20 Runs x 100 Prompts 6.000 93% (recommended) ~37 EUR
30 Runs x 200 Prompts 18.000 100% (Gold Standard) ~111 EUR

Rule of thumb: At least 1.500 observations per model (e.g. 30 Runs x 50 Prompts) for statistically reliable results. For small differences (under 5 percentage points) you need significantly more. More prompts can compensate for fewer runs.

Power >= 80% indicates a high probability of detecting a real difference. Below 60%, the risk of missing actual effects is too high.

What This Means for Brands (GEO Implications)

  1. Measure first, then optimize. You can't improve what you don't measure. This framework gives you a baseline.
  2. Model-specific strategies are essential. A brand that is visible on Gemini can be invisible on GPT.
  3. The prompt type makes a difference. Commercial prompts ("best brand for X") trigger different brands than informational prompts.
  4. Consistency matters. ESN has 20-36% across all models. AG1 has 5-62%. ESN has more stable AI visibility.
  5. Track over time. LLM training data changes. Monthly monitoring is recommended.

Interactive Report

You can view the complete interactive report with all charts, tables and raw data here:

Open Interactive HTML Report

All raw data (CSV) as download: Download result data (ZIP)

Open Source: Use It for Your Industry

The entire framework is open source and adaptable for any industry:

GitHub: github.com/AntonioBlago/llm-visibility-framework

  • 200 prompts (expandable), 48 brands (configurable), 3 models
  • Parallel API calls (ThreadPoolExecutor), automatic report generation
  • Interactive HTML report with Plotly charts
  • Statistics engine with power analysis
  • MIT license, free to use and customize

To continuously track and improve your AI visibility, check out Visibly AI, our SEO agent system with AI brand monitoring, competitor radar, and GEO optimization tools.

 
Cookie-Settings