Two structural reasons a single AI visibility check misleads, both visible in Google's numbers. (1) Google is not one AI surface but two - AI Overview and AI Mode - which behave like opposites: AI Mode is ~4x longer, names 2x more competitors, cites 1.7x more sources. (2) Neither surface is deterministic: ask the same question again and the brand-mention verdict flips on ~1 in 4 questions, and ~2 in 3 of the questions where the brand appears at all. Same instability on ChatGPT and Perplexity.

Everyone is talking about ChatGPT and Perplexity. Almost no one is talking about the largest AI search surface.
At Google I/O 2026, Google said AI Overviews now reach over 2.5 billion monthly users, and that AI Mode passed one billion monthly users barely a year after launch, with queries more than doubling every quarter.
AI Overviews alone is larger than ChatGPT, Claude and Perplexity combined, and AI Mode, past a billion users, is closing in.
So the GEO conversation is fixated on the wrong screen.
But the bigger problem is not which screen you watch, it is how you read it. Two structural blind spots make a single AI visibility check far less reliable than teams assume, and both are hiding inside Google's numbers.
Blind spot #1: Google is not one AI surface, it is two, and they behave like opposites. AI Overview (the snippet at the top of results) and AI Mode (the conversational search experience) name different numbers of brands, at different lengths, from different sources.
Other Articles
What Triggers ChatGPT Shopping, and What Happens Next?
Qwairy maps when ChatGPT Shopping appears across 100,000+ monitored runs. A matched test found appearance rates ranging from 21.7% to 96.7%.
ChatGPT Recommends the Product. Where Does the Shopping Link Go?
For 72 of 139 eligible product-brand domains, the recorded brand domain never appeared among 22,115 identifiable ChatGPT shopping destinations.
Blind spot #2: neither surface answers the same way twice. Ask the identical question again and the verdict on whether your brand is named flips a large share of the time. The same is true on ChatGPT and Perplexity.
We measured both across all seven major AI surfaces, on production brands we monitor continuously, with every prompt run repeatedly rather than checked once. The conclusion for anyone tracking AI visibility: a single check, on a single surface, captures neither the surface you are missing nor the noise in the one you are watching.
Google is two surfaces that behave like opposites. AI Mode answers run ~4x longer than AI Overview (median 2,944 vs 687 characters), name 2x more competitors per answer (6.16 vs 3.10), and cite 1.7x more sources (17.2 vs 10.0). One is a snippet; one is a chat answer.
AI Overview is the binary surface. It names 5.7x more competitors on a commercial query than on an explanatory one, and goes nearly silent on informational intent. AI Mode stays crowded across every intent.
Neither surface is deterministic. Ask the same question again and the brand-mention verdict flips on roughly 1 in 4 questions. Among the questions where a surface mentions your brand at all, it is inconsistent across runs about 2 in 3 of the time. We see the same instability on ChatGPT and Perplexity.
The differences between surfaces are real, not noise. In aggregate, over enough repeated runs, the run-to-run noise averages out and each surface settles into a stable, very different profile. That is exactly why you measure rates over many samples, not single answers.
We measure the content of the answer each surface returns for a monitored brand: how long it is, how many competing brands it names, which sources it cites, and whether it names the brand at all, across repeated runs. We do not measure how often an AI Overview appears on a Google search (an appearance rate), and we are explicit about why in the methodology.
Both of Google's AI surfaces, AI Overview and AI Mode, measured side by side on the same brands and an overlapping prompt set.
Five other assistants as a cross-surface control (ChatGPT, Perplexity, Gemini, Microsoft Copilot, Grok), the identical measurement on the same brands and window.
Period: April 1 to May 25, 2026 (~8 weeks).
Source: Qwairy Search Intelligence, production client brands only (audit and test brands excluded). Completed answers only. Crucially, each prompt is run repeatedly per surface rather than once, which is what lets us measure run-to-run stability, something a single-shot check cannot see.
This study measures the answer content captured by our platform: response length, competitor mentions, source citations, and per-run brand mention, read from the answer text itself. We report relative measures, rates, ratios and per-answer averages, rather than raw volume counts.
We deliberately do not report an AI Overview appearance rate ("X% of Google searches show an AI Overview"). A clean appearance rate requires logging every search performed, including those where no AI surface is triggered, so you have a denominator.
Our dataset records the AI answer when there is one to record; it is not built to count the searches that returned nothing. The numbers in this paper are all conditional on an answer existing.
We also do not headline per-query disagreement between surfaces as a structural finding. It is tempting (when at least one surface mentions the brand, two surfaces disagree on whether to name it ~45-48% of the time), but much of that per-query disagreement is the same run-to-run noise documented in Blind spot #2, not a stable property of the surfaces. The honest structural comparison is between aggregate rates, where the noise averages out. We report those.
The two Google surfaces diverge on every structural axis we measured.

Metric | AI Overview | AI Mode | Gap |
Median answer length | 687 chars (~110 words) | 2,944 chars (~470 words) | ~4.3x longer |
Competitor mention rate | 73.7% | 82.1% | +8.4 pp |
Avg competitors / answer | 3.10 |
The shape tells the story. AI Overview is a snippet: a short, top-of-results synthesis (median ~110 words) that names a handful of brands. AI Mode is a destination: a long-form, conversational answer (median ~470 words) that behaves much more like a chat assistant, naming twice as many brands and pulling in far more sources. They occupy the same Google search box in users' minds, but for a brand they are two different games. A brand can be visible in one and nearly absent from the other. One honest note on length: AI Mode's mean answer runs much longer than its median (a minority of very long answers pulls the average up past 10,000 characters). We report the median throughout because it is the honest center of the distribution.
Run a free audit: see if ChatGPT, Gemini and Copilot recommend you, in about a minute.
Here is the finding most likely to change how you read any AI visibility number. Because we run each prompt multiple times per surface, we can ask a simple question: when a surface mentions your brand on one run, does it mention it on the next? Often, no.

Surface | Questions where the verdict flips across runs | Of questions where the brand ever appears, share inconsistent across runs |
AI Overview | 24.7% | 69.5% |
AI Mode | 24.7% | 65.3% |
OpenAI ChatGPT | 21.8% | 66.0% |
Perplexity |
On roughly 1 in 4 questions, a surface gives an inconsistent verdict on your brand across its own runs: present in some, absent in others. And when you narrow to the questions where the brand appears at all, the inconsistency is the rule, not the exception: about 2 in 3 of those questions are unstable across runs. This holds on every surface we tested, Google and non-Google alike. The practical consequence is blunt.
A single check is close to a coin flip. "I asked ChatGPT and it didn't mention us" is not evidence of invisibility; it is one draw from a distribution. The only reliable read is a rate measured over many runs, which is what makes the structural differences in Part 1 trustworthy in the first place: individual answers are noisy, so you average enough of them until the signal is stable. That is the entire case for systematic, repeated measurement over one-off prompting.
If individual answers are this noisy, how do we know the AI Overview vs AI Mode gap in Part 1 is real and not more noise? Because aggregate rates over enough repeated runs are stable, and we can see each surface settle into a distinct profile. The cross-surface control makes this visible: run the identical measurement on five other assistants, on the same brands and window, and each lands at its own stable level.

Google's two surfaces sit at opposite ends of the field: AI Overview is the second-most-selective surface we measure, AI Mode among the most crowded, with three other assistants in the gap between them. This is also the artifact check: if the AI Overview vs AI Mode gap were a quirk of how we collect each surface, five independent assistants would not slot neatly in between. They do. "Google" is not a setting; it is two products with different jobs, and the gap is structural.
Both Google surfaces name more brands on commercial-intent questions than on informational ones, the universal buyer-journey pattern. But AI Overview discriminates far more sharply.

On AI Overview, a comparison query surfaces 5.7x more competitors than an explanatory one (4.62 vs 0.81), and on explanatory questions it mentions a competitor less than half the time. On informational intent, AI Overview is effectively a no-show for competitive positioning.
AI Mode never drops below 2.36 competitors per answer, even on explanations. On AI Overview, your competitive battleground is almost entirely commercial-intent queries; on AI Mode, you are in a crowded room on nearly every query.
See your mentions across ChatGPT, Claude and Perplexity in real time, the moment buyers ask.
Never trust a single check. Visibility is a rate, not a yes/no. The same surface flips its verdict on ~1 in 4 questions and is inconsistent ~2 in 3 of the time when your brand is in play. Measure over many runs and over time, or you are reading noise.
Track Google as two surfaces, and track more than one engine. A brand can win AI Mode and lose AI Overview, or the reverse, and no engine is a proxy for another. Blending them, or watching only ChatGPT, hides real movement.
For AI Overview, win commercial intent. It is the snippet 2.5 billion people see, and your competitive presence there is concentrated on comparisons, "best X" lists, and recommendations. Informational content has little competitive leverage on this surface.
For AI Mode, expect a crowded room everywhere. With twice the competitors per answer and 470-word responses, mention alone is cheap; position and differentiation are what matter.
Reality: one run is a single draw. The same surface is inconsistent across runs ~2 in 3 of the time when the brand is in play. Measure the rate over many runs before concluding anything.
Reality: AI Overview and AI Mode diverge on length (~4x), competitor density (~2x), and sourcing (~1.7x). One playbook cannot serve both.
Reality: AI Overview names 5.7x more competitors on commercial intent than on explanatory intent, and barely mentions competitors on explanations at all. For competitive visibility, commercial-intent content is the lever.
Reality: AI Mode mentions a competitor in 82% of answers and names 6+ per answer. Mention is the floor, not the win. Position and differentiation decide outcomes.
Scope: Google's two AI surfaces (AI Overview, AI Mode) plus five other assistants (ChatGPT, Perplexity, Gemini, Microsoft Copilot, Grok) as a cross-surface control, all measured on production client brands over April 1 to May 25, 2026 (audit and test brands excluded). Every prompt is run repeatedly per surface, which is what lets us measure run-to-run stability.
Methodology: Quantitative analysis of completed AI answers. Response length measured in characters from the answer text (medians reported). Competitor mentions and source citations measured per answer and aggregated per surface. Run-to-run stability measured on prompts with two or more runs per surface in the window: a "flip" means the brand is named on some runs and not others. Intent split uses the answer-type classification (comparison, list, recommendation, how-to, explanation). No appearance-rate metric is computed (see Limitations).
Non-determinism is measured, not eliminated. We quantify run-to-run instability (a surface flips its brand verdict on ~1 in 4 questions; ~2 in 3 conditional on ever-mentioning). Some of this is genuine model stochasticity and some is retrieval variance over time; we do not separate the two. The takeaway (single checks are unreliable, rates are not) holds either way.
No appearance rate. This dataset is not built to measure how often an AI Overview appears on a Google search. Every figure here is conditional on an answer existing.
Per-query cross-surface disagreement is not reported as structural. It is dominated by the same run-to-run noise above. We compare aggregate rates instead, where noise averages out.
Monitored-brand population. These are brands actively tracked by their owners, not a neutral sample. The cross-surface comparisons control for this because all surfaces see the same brands.
Snapshot. Both Google surfaces are evolving fast (AI Mode queries are doubling quarterly per Google). This is an 8-week snapshot; a repeat in Q3 2026 may diverge.
All figures are conditional on an answer existing; no AI Overview appearance rate is claimed.
We report relative measures (rates, ratios, per-answer averages), not raw volume counts.
We do not report per-query disagreement between surfaces as a structural finding, because it is confounded by run-to-run instability. We compare aggregate rates instead.
Google I/O 2026, Search updates (AI Mode passed 1B monthly users, queries doubling every quarter)
Sundar Pichai: AI Overviews now has over 2.5 billion monthly users (CNBC, May 19, 2026)
Related research: LLM Behavior Study Q1 2026, Sources by Intent Study Q2 2026
New to the topic? Start with What is GEO?
Want to track your brand across both Google AI surfaces, with repeated sampling instead of one-off checks? Qwairy monitors brand mentions, position and citations across AI Overview, AI Mode, ChatGPT, Perplexity, Claude, Gemini and expanding providers, with per-surface analytics.
Track your mentions across ChatGPT, Claude, Perplexity and all major AI platforms. Join 1,500+ brands monitoring their AI presence in real-time.
Free trial • No credit card required • Complete platform access
~2.0x more |
Avg sources / answer | 9.99 | 17.20 | ~1.7x more |
Answers with a source | 96.3% | 97.6% | both near-universal |
65.5% |
Surface | Mention rate | Avg competitors / answer | Avg sources / answer |
Perplexity | 63.9% | 2.86 | 9.61 |
Google AI Overview | 73.7% | 3.10 | 9.99 |
OpenAI ChatGPT | 72.8% | 5.17 | 4.16 |
Google Gemini | 78.6% | 6.13 | 1.31 |
Google AI Mode | 82.1% | 6.16 | 17.20 |
Microsoft Copilot | 92.7% | 6.99 | 3.51 |
xAI Grok | 94.5% | 12.68 | 43.34 |
Question type | AI Overview avg competitors | AI Mode avg competitors |
Comparison | 4.62 | 7.39 |
List ("best X") | 4.11 | 10.36 |
Recommendation | 3.69 | 8.06 |
How-to | 1.63 | 4.33 |
Explanation | 0.81 | 2.36 |