How to measure AI models for brand visibility

Discover how to measure AI models reliably so you can build a Generative Engine Optimization (GEO) campaign that delivers results

Let's use the example of running shoe prompts.Sampled once gives you a high-level score; sampled 100 times, you see where Adidas wins and loses.

Other GEO platforms

100 unique prompts each asked once

29% visibility, ±9 pt margin of error

100 unique prompts each asked 100 times

29% visibility, ±1 pt margin of error

= one response = Adidas is mentioned

Source: Evertune analysis of 10,000 prompt responses on running shoes, July 2026

Why does sampling each prompt 100 times lead to statistical significance?

Let's use an example. Is Adidas mentioned when you ask AI models once: "What are the best cushioned shoes for marathons?"

Each dot = one response to this prompt.
Blue dots = Adidas is mentioned.

Visibility: Mentioned · Margin of error: ± 44 points

Asked once and Adidas is mentioned in the response, leading you to presume Adidas' AI visibility is 100%.

Ask 25 times and the score becomes 36%, however the margin of error is ±20 points.

Ask 100 times and you can trust the result: Adidas' visibility on this prompt is about 48%, with a margin of error of ±10 points.

Why not just sample many unique prompts once?

Many GEO providers argue you don't need to sample the same prompts often. Ask across a broad enough set of prompts, they say, and it all averages out. It doesn't, and here's why.

Let's look at 100 different prompts about running shoes, each sampled once

Each dot = one response to prompts about running shoes, each asked once.
Blue dots = Adidas is mentioned.

Adidas' overall visibility: 31%

These 100 prompts all explore one category, running shoes, across different topics like budget picks, beginner runners, and marathon training. When each prompt is asked once, Adidas' overall visibility scores 29% (±9 points): seemingly stable enough to trust at the brand level.

But only 10 of those 100 prompts dig into one topic, such as "budget picks." Because each was asked once, Adidas' budget score rests on just 10 responses, so its margin of error balloons to ±31 points. A range that large could mean that Adidas' visibility score could be as low as 19% or as high as 81% on budget-related prompts, telling vastly different stories.

Sample each of the 100 prompts 100 times and every prompt becomes its own stable measurement. Instead of 100 responses, you have 10,000, and the scores stop jumping from run to run.

Cluster the 100 prompts into 10 topics of 10 prompts each, and every topic now carries 1,000 responses. That tightens the margin of error to ±1 point on overall visibility and ±2 points at the topic level. Precise enough to rank topics, compare them and invest accordingly.

Read the full research report

What this means for Adidas

Now Adidas knows exactly where they stand: strong on race-day, weak on trail running. That tells them where to invest: reinforce what's working, fix what the models get wrong.

Grounded in real buyer behavior

The questions you track decide whether your data reflects reality.

Every prompt Evertune measures comes from EverPanel, our panel of over 150 million real conversations, weighted to reflect the actual composition of the internet. You're measured on the language and questions buyers actually use in your category, not a handful of cherry-picked queries. And because over 80% of AI questions are phrased uniquely, Evertune groups prompts by intent and topic, so you see the full picture of how buyers explore your category.

See what buyers ask in your category

The questions buyers ask before they name any brand are the consideration moments you can still win.

See how buyers ask about you

Where buyers ask AI about you directly: comparing you, evaluating you or trying to reach you.

See how often they're asking

How demand for any brand, product or topic is trending across every major model, so you focus where the volume is.

Before and during live search

What AI already knows, and what AI finds in the moment.

When AI answers a prompt, it draws on what it already knows about your brand and what it looks up in the moment. Most tools measure only the second. Evertune measures both.

What AI already knows

Through direct API access, Evertune isolates the model's foundational knowledge: the view of your brand baked in during training, including the preferences and biases it holds before any search happens.

What AI finds in the moment

The live consumer apps answer buyers by searching the web as the question is asked and citing what they find. This is what your buyers see today, and it shifts as you optimize for AI search.

The gap tells you what to fix

Strong in foundational knowledge but weak in the consumer app points to an SEO and content gap. Weak in both points to a brand-building gap. The difference tells you where to invest.

Book a demo

Speak with our team to see how Evertune can help you own the AI customer journey.

Request a demo