Anyone can screenshot a chatbot. A measurement needs a denominator, a method and a change log. This page is all three.
Autumn 2026 edition · data collected 3 to 7 September 2026 · 178,786 answers
Written and maintained by Matias Garrido-Garcia, founder. Corrections and questions: write to us.
We track 185 supplement shelves: product categories like magnesium glycinate, prenatal multivitamin, lion's mane or collagen. Each shelf is sampled with a battery of shopping prompts, the questions a real buyer types, repeated hundreds of times on every AI assistant. A few shelves are a shopping goal rather than a single ingredient, such as immune support; we keep them because shoppers ask for them by that name.
Every question is hand-curated per shelf: a human decides what a real buyer would ask, in buyer language. That is only feasible because AI Shelf Space covers one industry; a tool serving every industry cannot hand-curate every question.
AI answers are probabilistic. In an independent study of roughly 3,000 prompt runs, fewer than 1% of identical prompts returned identical brand lists, and it took on the order of 60 to 100 runs for the brand set to stabilize. A tool that asks once is reporting a coin flip.
AI Shelf Space samples every shelf hundreds of times on every AI assistant, every edition, and the number of answers behind every figure is shown next to it. An answer that names no brand counts as recommending no brand. Visibility is computed on the answers that named at least one brand, and that share is published per assistant below, so the denominator is never hidden.
Reading our numbers: "Visibility 76% on ChatGPT" means: of the answers on that shelf that recommended any brand this edition, 76% recommended this one. The assistant's brand-carrying rate is always published next to it.
Why your own test will look different. One chat is one draw. Two independent batches of the same question, run on the same day, moved a brand's share by 16 points on average (8 points on ChatGPT, 22 on Gemini). A dashboard built on hundreds of answers shows the weight of the coin, not one toss, which is also why a rank is the more robust figure: it cannot be contradicted by a single chat.
We currently sample ChatGPT, Google Gemini, Google AI Overviews, Google AI Mode, Microsoft Copilot, Perplexity and Claude. Consumer apps are sampled as shoppers see them, not through developer APIs. Claude is the exception: no vendor exposes its consumer UI at scale, so it is sampled via Anthropic's API with web search enabled, and flagged accordingly. Model versions are re-verified every edition and pinned in the change log (§8).
Share of answers naming at least one brand, per assistant, this edition:
| Assistant | Brand-carrying rate | Collection |
|---|---|---|
| Microsoft Copilot | 99.6% | consumer app |
| Google AI Mode | 97.5% | consumer app |
| Google AI Overviews | 96.0% | consumer app |
| Perplexity | 95.6% | consumer app |
| Google Gemini | 95.2% | consumer app |
| Claude | 91.6% | API + web search |
| ChatGPT | 90.5% | consumer app |
Each raw answer is parsed by an extraction model into brands, products and list positions. Brand names are then canonicalized: "GOL," "Garden of Life" and "gardenoflife.com" count as one brand, matched against an alias map covering 9,857 brand names this edition. Every data point carries the brand-map version that produced it, so results are reproducible.
The public Index publishes a new edition every quarter, each built on a full measurement of every shelf on every assistant. Between editions, customers' shelves are re-sampled weekly for the alerts. Every edition states the week its data was collected, and a published edition is never silently rewritten.