Product updateQwairy v1.18
Qwairy v1.18: AI Revenue, Action Center & Pitch AuditRead the article
Qwairy
  • Pricing
  • Agencies
  • Blog
    116
Log inGet a demoStart Free
Qwairy

Optimize your visibility in the AI era with advanced Generative Engine Optimization.

Platform

  • Cockpit
  • Monitor
  • Act
  • Analyze
  • Optimize
  • Measure

Product

  • Pricing
  • Integrations
  • Documentation
  • API/MCP
  • Changelog
  • AffiliatesNew

Solutions

  • For Teams
  • Compare

Resources

  • Free AI Visibility AuditNew
  • Blog
  • GEO Guide
  • GEO Glossary
  • AI Crawlers Guide
  • Help Center

Company

  • About Us
  • Get a demo
  • Privacy Policy
  • Terms of Service
  • Legal Notice

© 2026 Qwairy SAS. All rights reserved.

GDPR Compliant
🇪🇺EU Data Hosting

Made with ❤️ in France 🇫🇷

  1. Home/
  2. Blog/
  3. AI Visibility Blind Spots May 2026
GEO
AI Visibility
ChatGPT
AI Search
2026

Same Question, Different Answer: The Two Blind Spots in How You Track AI Visibility (May 2026)

Two structural reasons a single AI visibility check misleads, both visible in Google's numbers. (1) Google is not one AI surface but two - AI Overview and AI Mode - which behave like opposites: AI Mode is ~4x longer, names 2x more competitors, cites 1.7x more sources. (2) Neither surface is deterministic: ask the same question again and the brand-mention verdict flips on ~1 in 4 questions, and ~2 in 3 of the questions where the brand appears at all. Same instability on ChatGPT and Perplexity.

Qwairy•May 29, 2026•15 min read•
Research
Summarize with AI

Everyone is talking about ChatGPT and Perplexity. Almost no one is talking about the largest AI search surface.

At Google I/O 2026, Google said AI Overviews now reach over 2.5 billion monthly users, and that AI Mode passed one billion monthly users barely a year after launch, with queries more than doubling every quarter.

AI Overviews alone is larger than ChatGPT, Claude and Perplexity combined, and AI Mode, past a billion users, is closing in.

So the GEO conversation is fixated on the wrong screen.

But the bigger problem is not which screen you watch, it is how you read it. Two structural blind spots make a single AI visibility check far less reliable than teams assume, and both are hiding inside Google's numbers.

Blind spot #1: Google is not one AI surface, it is two, and they behave like opposites. AI Overview (the snippet at the top of results) and AI Mode (the conversational search experience) name different numbers of brands, at different lengths, from different sources.

Other Articles

What Triggers ChatGPT Shopping, and What Happens Next?

Qwairy maps when ChatGPT Shopping appears across 100,000+ monitored runs. A matched test found appearance rates ranging from 21.7% to 96.7%.

8/5/2026•6 min read

ChatGPT Recommends the Product. Where Does the Shopping Link Go?

For 72 of 139 eligible product-brand domains, the recorded brand domain never appeared among 22,115 identifiable ChatGPT shopping destinations.

8/5/2026•6 min read
View all articles

Blind spot #2: neither surface answers the same way twice. Ask the identical question again and the verdict on whether your brand is named flips a large share of the time. The same is true on ChatGPT and Perplexity.

We measured both across all seven major AI surfaces, on production brands we monitor continuously, with every prompt run repeatedly rather than checked once. The conclusion for anyone tracking AI visibility: a single check, on a single surface, captures neither the surface you are missing nor the noise in the one you are watching.

What we found

  • Google is two surfaces that behave like opposites. AI Mode answers run ~4x longer than AI Overview (median 2,944 vs 687 characters), name 2x more competitors per answer (6.16 vs 3.10), and cite 1.7x more sources (17.2 vs 10.0). One is a snippet; one is a chat answer.

  • AI Overview is the binary surface. It names 5.7x more competitors on a commercial query than on an explanatory one, and goes nearly silent on informational intent. AI Mode stays crowded across every intent.

  • Neither surface is deterministic. Ask the same question again and the brand-mention verdict flips on roughly 1 in 4 questions. Among the questions where a surface mentions your brand at all, it is inconsistent across runs about 2 in 3 of the time. We see the same instability on ChatGPT and Perplexity.

  • The differences between surfaces are real, not noise. In aggregate, over enough repeated runs, the run-to-run noise averages out and each surface settles into a stable, very different profile. That is exactly why you measure rates over many samples, not single answers.

We measure the content of the answer each surface returns for a monitored brand: how long it is, how many competing brands it names, which sources it cites, and whether it names the brand at all, across repeated runs. We do not measure how often an AI Overview appears on a Google search (an appearance rate), and we are explicit about why in the methodology.

Data and methodology

Scope

  • Both of Google's AI surfaces, AI Overview and AI Mode, measured side by side on the same brands and an overlapping prompt set.

  • Five other assistants as a cross-surface control (ChatGPT, Perplexity, Gemini, Microsoft Copilot, Grok), the identical measurement on the same brands and window.

  • Period: April 1 to May 25, 2026 (~8 weeks).

  • Source: Qwairy Search Intelligence, production client brands only (audit and test brands excluded). Completed answers only. Crucially, each prompt is run repeatedly per surface rather than once, which is what lets us measure run-to-run stability, something a single-shot check cannot see.

What we measure, and what we deliberately do not

This study measures the answer content captured by our platform: response length, competitor mentions, source citations, and per-run brand mention, read from the answer text itself. We report relative measures, rates, ratios and per-answer averages, rather than raw volume counts.

We deliberately do not report an AI Overview appearance rate ("X% of Google searches show an AI Overview"). A clean appearance rate requires logging every search performed, including those where no AI surface is triggered, so you have a denominator.

Our dataset records the AI answer when there is one to record; it is not built to count the searches that returned nothing. The numbers in this paper are all conditional on an answer existing.

We also do not headline per-query disagreement between surfaces as a structural finding. It is tempting (when at least one surface mentions the brand, two surfaces disagree on whether to name it ~45-48% of the time), but much of that per-query disagreement is the same run-to-run noise documented in Blind spot #2, not a stable property of the surfaces. The honest structural comparison is between aggregate rates, where the noise averages out. We report those.

Part 1, two surfaces, opposite shapes

The two Google surfaces diverge on every structural axis we measured.

AI Overview and AI Mode behave like opposites: median length, competitors per answer, and sources per answer all far higher on AI Mode
Metric
AI Overview
AI Mode
Gap
Median answer length
687 chars (~110 words)
2,944 chars (~470 words)
~4.3x longer
Competitor mention rate
73.7%
82.1%
+8.4 pp
Avg competitors / answer
3.10

The shape tells the story. AI Overview is a snippet: a short, top-of-results synthesis (median ~110 words) that names a handful of brands. AI Mode is a destination: a long-form, conversational answer (median ~470 words) that behaves much more like a chat assistant, naming twice as many brands and pulling in far more sources. They occupy the same Google search box in users' minds, but for a brand they are two different games. A brand can be visible in one and nearly absent from the other. One honest note on length: AI Mode's mean answer runs much longer than its median (a minority of very long answers pulls the average up past 10,000 characters). We report the median throughout because it is the honest center of the distribution.

Is your brand visible in AI search?

Run a free audit: see if ChatGPT, Gemini and Copilot recommend you, in about a minute.

Run my free audit

Part 2, neither surface answers the same way twice

Here is the finding most likely to change how you read any AI visibility number. Because we run each prompt multiple times per surface, we can ask a simple question: when a surface mentions your brand on one run, does it mention it on the next? Often, no.

Same surface, same question, different answer: across repeated runs, each surface fails to mention a brand on at least one run roughly two thirds of the time it mentions it at all
Surface
Questions where the verdict flips across runs
Of questions where the brand ever appears, share inconsistent across runs
AI Overview
24.7%
69.5%
AI Mode
24.7%
65.3%
OpenAI ChatGPT
21.8%
66.0%
Perplexity

On roughly 1 in 4 questions, a surface gives an inconsistent verdict on your brand across its own runs: present in some, absent in others. And when you narrow to the questions where the brand appears at all, the inconsistency is the rule, not the exception: about 2 in 3 of those questions are unstable across runs. This holds on every surface we tested, Google and non-Google alike. The practical consequence is blunt.

A single check is close to a coin flip. "I asked ChatGPT and it didn't mention us" is not evidence of invisibility; it is one draw from a distribution. The only reliable read is a rate measured over many runs, which is what makes the structural differences in Part 1 trustworthy in the first place: individual answers are noisy, so you average enough of them until the signal is stable. That is the entire case for systematic, repeated measurement over one-off prompting.

Part 3, the differences between surfaces are real, not noise

If individual answers are this noisy, how do we know the AI Overview vs AI Mode gap in Part 1 is real and not more noise? Because aggregate rates over enough repeated runs are stable, and we can see each surface settle into a distinct profile. The cross-surface control makes this visible: run the identical measurement on five other assistants, on the same brands and window, and each lands at its own stable level.

Average competitors per answer across seven AI surfaces: Google's AI Overview sits near the bottom and AI Mode near the top, with the other assistants in between

Google's two surfaces sit at opposite ends of the field: AI Overview is the second-most-selective surface we measure, AI Mode among the most crowded, with three other assistants in the gap between them. This is also the artifact check: if the AI Overview vs AI Mode gap were a quirk of how we collect each surface, five independent assistants would not slot neatly in between. They do. "Google" is not a setting; it is two products with different jobs, and the gap is structural.

Part 4, AI Overview is all-in or silent

Both Google surfaces name more brands on commercial-intent questions than on informational ones, the universal buyer-journey pattern. But AI Overview discriminates far more sharply.

Average competitors per answer by question type for AI Overview vs AI Mode: AI Overview collapses to near zero on explanatory questions while AI Mode stays high

On AI Overview, a comparison query surfaces 5.7x more competitors than an explanatory one (4.62 vs 0.81), and on explanatory questions it mentions a competitor less than half the time. On informational intent, AI Overview is effectively a no-show for competitive positioning.

AI Mode never drops below 2.36 competitors per answer, even on explanations. On AI Overview, your competitive battleground is almost entirely commercial-intent queries; on AI Mode, you are in a crowded room on nearly every query.

Is your brand visible in AI search?

See your mentions across ChatGPT, Claude and Perplexity in real time, the moment buyers ask.

Check now

What this means for your GEO strategy

Never trust a single check. Visibility is a rate, not a yes/no. The same surface flips its verdict on ~1 in 4 questions and is inconsistent ~2 in 3 of the time when your brand is in play. Measure over many runs and over time, or you are reading noise.

Track Google as two surfaces, and track more than one engine. A brand can win AI Mode and lose AI Overview, or the reverse, and no engine is a proxy for another. Blending them, or watching only ChatGPT, hides real movement.

For AI Overview, win commercial intent. It is the snippet 2.5 billion people see, and your competitive presence there is concentrated on comparisons, "best X" lists, and recommendations. Informational content has little competitive leverage on this surface.

For AI Mode, expect a crowded room everywhere. With twice the competitors per answer and 470-word responses, mention alone is cheap; position and differentiation are what matter.

Common mistakes

"I asked the AI once and it didn't mention us, so we're invisible"

Reality: one run is a single draw. The same surface is inconsistent across runs ~2 in 3 of the time when the brand is in play. Measure the rate over many runs before concluding anything.

"Google AI is one thing, optimize for it once"

Reality: AI Overview and AI Mode diverge on length (~4x), competitor density (~2x), and sourcing (~1.7x). One playbook cannot serve both.

"AI Overview is informational, so educational content wins there"

Reality: AI Overview names 5.7x more competitors on commercial intent than on explanatory intent, and barely mentions competitors on explanations at all. For competitive visibility, commercial-intent content is the lever.

"Getting mentioned in AI Mode is the goal"

Reality: AI Mode mentions a competitor in 82% of answers and names 6+ per answer. Mention is the floor, not the win. Position and differentiation decide outcomes.

About this study

Scope: Google's two AI surfaces (AI Overview, AI Mode) plus five other assistants (ChatGPT, Perplexity, Gemini, Microsoft Copilot, Grok) as a cross-surface control, all measured on production client brands over April 1 to May 25, 2026 (audit and test brands excluded). Every prompt is run repeatedly per surface, which is what lets us measure run-to-run stability.

Methodology: Quantitative analysis of completed AI answers. Response length measured in characters from the answer text (medians reported). Competitor mentions and source citations measured per answer and aggregated per surface. Run-to-run stability measured on prompts with two or more runs per surface in the window: a "flip" means the brand is named on some runs and not others. Intent split uses the answer-type classification (comparison, list, recommendation, how-to, explanation). No appearance-rate metric is computed (see Limitations).

Limitations

Non-determinism is measured, not eliminated. We quantify run-to-run instability (a surface flips its brand verdict on ~1 in 4 questions; ~2 in 3 conditional on ever-mentioning). Some of this is genuine model stochasticity and some is retrieval variance over time; we do not separate the two. The takeaway (single checks are unreliable, rates are not) holds either way.

No appearance rate. This dataset is not built to measure how often an AI Overview appears on a Google search. Every figure here is conditional on an answer existing.

Per-query cross-surface disagreement is not reported as structural. It is dominated by the same run-to-run noise above. We compare aggregate rates instead, where noise averages out.

Monitored-brand population. These are brands actively tracked by their owners, not a neutral sample. The cross-surface comparisons control for this because all surfaces see the same brands.

Snapshot. Both Google surfaces are evolving fast (AI Mode queries are doubling quarterly per Google). This is an 8-week snapshot; a repeat in Q3 2026 may diverge.

Transparency notes

  • All figures are conditional on an answer existing; no AI Overview appearance rate is claimed.

  • We report relative measures (rates, ratios, per-answer averages), not raw volume counts.

  • We do not report per-query disagreement between surfaces as a structural finding, because it is confounded by run-to-run instability. We compare aggregate rates instead.

Sources and references

  • Google I/O 2026, Search updates (AI Mode passed 1B monthly users, queries doubling every quarter)

  • Sundar Pichai: AI Overviews now has over 2.5 billion monthly users (CNBC, May 19, 2026)

  • Related research: LLM Behavior Study Q1 2026, Sources by Intent Study Q2 2026

  • New to the topic? Start with What is GEO?

Want to track your brand across both Google AI surfaces, with repeated sampling instead of one-off checks? Qwairy monitors brand mentions, position and citations across AI Overview, AI Mode, ChatGPT, Perplexity, Claude, Gemini and expanding providers, with per-surface analytics.

Start Monitoring Today

Is Your Brand Visible in AI Search?

Track your mentions across ChatGPT, Claude, Perplexity and all major AI platforms. Join 1,500+ brands monitoring their AI presence in real-time.

Complete AI Monitoring
Track every mention in real-time
Competitor Intelligence
See what AI recommends
Proven Results
87% see improvements in 30 days
Start Free Trial

Free trial • No credit card required • Complete platform access

In this article

  • What we found
  • Data and methodology
  • Part 1, two surfaces, opposite shapes
  • Part 2, neither surface answers the same way twice
  • Part 3, the differences between surfaces are real, not noise
  • Part 4, AI Overview is all-in or silent
  • What this means for your GEO strategy
  • Common mistakes
  • About this study

Share

See your brand in AI search

Book a demo and discover how you rank across ChatGPT, Claude and Perplexity.

Book a demo
6.16
~2.0x more
Avg sources / answer
9.99
17.20
~1.7x more
Answers with a source
96.3%
97.6%
both near-universal
20.6%
65.5%
Surface
Mention rate
Avg competitors / answer
Avg sources / answer
Perplexity
63.9%
2.86
9.61
Google AI Overview
73.7%
3.10
9.99
OpenAI ChatGPT
72.8%
5.17
4.16
Google Gemini
78.6%
6.13
1.31
Google AI Mode
82.1%
6.16
17.20
Microsoft Copilot
92.7%
6.99
3.51
xAI Grok
94.5%
12.68
43.34
Question type
AI Overview avg competitors
AI Mode avg competitors
Comparison
4.62
7.39
List ("best X")
4.11
10.36
Recommendation
3.69
8.06
How-to
1.63
4.33
Explanation
0.81
2.36
FAQ
What are the two blind spots in AI visibility tracking? First, Google is two AI surfaces, not one: AI Overview (the snippet) and AI Mode (the conversational experience) behave like opposites, so a brand visible in one can be absent from the other. Second, neither surface is deterministic: ask the same question again and the brand-mention verdict flips on ~1 in 4 questions, and ~2 in 3 of the questions where the brand appears at all. A single check on a single surface misses both.
What is the difference between Google AI Overview and AI Mode? AI Overview is the short AI snippet at the top of a normal Google results page (median ~110 words in our data). AI Mode is Google's conversational, chat-style search surface that returns long-form answers (median ~470 words). On our data AI Mode names about twice as many competitors per answer and cites about 1.7x more sources.
If AI answers are this noisy, how can the surface differences be trusted? Because noise averages out. Individual answers are unstable run to run, but aggregate rates over enough repeated runs are stable, and each surface settles into a distinct profile. That is precisely why visibility should be measured as a rate over many samples, not read off a single answer.
Why not report how often an AI Overview appears on a search? A clean appearance rate needs a denominator: every search run, including those that return no AI surface. Our dataset records the answer when there is one to record; it is not built to count searches that returned nothing. Every figure here is conditional on an answer existing.
Which Google surface should I optimize for first? Both, separately. AI Overview reaches 2.5 billion monthly users and rewards commercial-intent content (comparisons, "best X", recommendations). AI Mode is growing fastest and is crowded on every intent, so position and differentiation matter more than mention. Treat them as two channels, and measure each with repeated sampling.