Which AI Engines Can You Actually Measure? One Reports Citations, One Reports Impressions, Three Report Nothing.
Every citation number I publish comes from one engine family: Microsoft Copilot and its partner surfaces. For months I described that as the whole story, as though one engine were measurable and the other four were dark. The framing was too clean, and it stopped being true seven weeks before I noticed.
The real shape has three tiers. Which tier an engine sits in decides what anybody can honestly claim about your visibility on it, including me. Checked 28 July 2026, because three of these five entries changed inside the last year.
Tier one: citations, with the queries attached
Microsoft Copilot and its partner surfaces. Two free tools report back to you. Bing Webmaster Tools has an AI Performance report that counts how many times your pages were cited, day by day. Microsoft Clarity has an AI citations dashboard that adds the grounding queries and your share of the cited sources on them. Both are free, both export, and both are first-party, meaning Microsoft is telling you what Microsoft did.
Worth naming the incentive out loud: Microsoft looks better when publishers believe Copilot sends value. I hold that number to the same suspicion I would hold anyone else's, which is why I spent a weekend counting it by hand.
This tier has one member. That part has not changed.
Tier two: impressions, if your property has the report
Google. In June 2026 Google began rolling out generative AI performance reporting inside Search Console. It shows impressions from Google's AI surfaces, which tells you that you appeared without telling you what you were quoted for.
Two things to hold at once here. It is a staged rollout and plenty of properties do not have it, my mom's shop included, so check yours before you assume you can measure this. And an impression is a weaker fact than a citation. Knowing you showed up is useful. Knowing which question you showed up on, and who got named beside you, is what you plan against.
I got this one wrong in print. I wrote that no Google AI report existed, seven weeks after Google shipped one. That is not a small error, and the correction is why every claim on this page carries a date.
Separate control, constantly confused with this one: Google-Extended governs Gemini apps and Vertex AI grounding, and Google's documentation states it does not affect Search inclusion. AI Overviews live inside Search, so the levers there are the ordinary snippet directives.
Tier three: nothing at all
ChatGPT, Claude and Perplexity. No owner-facing report of any kind.
Server logs get you closer than most people assume, as long as you read them carefully, because these companies run separate agents for separate jobs and the names are not interchangeable. OpenAI runs GPTBot for training, OAI-SearchBot for appearing in ChatGPT's search results, and ChatGPT-User for the live fetch made while answering a question. Anthropic splits the same way, with ClaudeBot, Claude-User and Claude-SearchBot. Perplexity runs PerplexityBot, plus Perplexity-User, which its own documentation says generally ignores robots.txt because a person asked for that specific fetch.
A hit from a fetch agent is a warmer signal than a training crawl and worth separating in your logs. What no log tells you is whether the fetched page ended up in the answer.
Perplexity is the friendliest of the three to audit by hand, because it prints its sources on nearly every answer. You can count. You just have to count it yourself.
What the tiers mean when somebody sells you a number
When a vendor shows you a chart of your visibility across all the major AI engines, the tier tells you what you are looking at before you read a single figure. Anything on the Copilot row can be exported and checked. The Google row is an impression at best, and may not exist for your property at all. The other three rows were sampled, because sampling is the only method there is on them.
Sampling is a legitimate method. I use it. It stops being legitimate the moment somebody prints the output next to first-party numbers without saying which is which. So ask two questions before you sign anything: what instrument produced this, and how many runs sit behind it.
So I checked the one engine that reports citations, by hand
A dashboard is a claim. I wanted to know whether the Copilot number I keep publishing survives being checked independently, so I reproduced it manually.
Twenty runs, using the grounding queries Clarity said my mom's shop already wins, asked one at a time, with every source slot in every answer logged. 192 source slots in total. I counted how many belonged to her.
61 of 192, which is 31.8%. Clarity reported 33.2% for the same period.
Before anyone quotes that as proof: 61 events carries a margin of roughly seven points, and my own run set swings far wider than that inside a single query. A 1.4-point gap sits comfortably inside the noise. The queries were also ones Clarity told me the shop already wins, so I checked the dashboard against a set the dashboard chose. That is a consistency check, not an independent validation, and I would not let a vendor call it one.
What I can honestly say is that the dashboard did not turn out to be wildly wrong when I counted by hand, on the queries it nominated. That is worth something and it is less than a validation.
And then the part nobody publishes
While counting, I found something that matters more than the agreement.
The same query, asked twice in a row, scored 0% and then 69%.
Not a different question and not a different day, but the identical string asked back to back. Across the run set, per-cell standard deviations ran as high as 49 points.
I read that as generative engines sampling rather than looking up, though caching, experiment bucketing and rate limiting on rapid repeats would produce the same observation and I have not ruled them out. Whatever the mechanism, the consequence for you is identical: two people asking your industry's most important question ten seconds apart can get answers with completely different source lists.
Here is what follows, and it is the practical part: one run is not a measurement, it is a coin flip you wrote down. Any AI-visibility screenshot where somebody asked an engine a question once is a single coin flip presented as a finding.
The mistake I nearly published
I also ran an experiment. I rewrote some product pages in the style of the one page that earns most of the shop's citations, to see whether the phrasing itself was doing the work.
Pooled across all the runs, the reworded pages looked better by 34 points, with a confidence interval that excluded zero. It looked like a finding. I had the sentence half written.
Then I checked how the queries were distributed between the two arms, and they were not matched. The reworded arm carried three runs of a query with no control, and the other arm carried the single query that scored 100%. The comparison was measuring which questions landed in which group, not the wording.
Paired properly, on the four queries that appeared in both arms, the difference was not significant. One of the four ran backwards. My power calculation said detecting a 20-point effect at this variance would need roughly 113 matched queries, which is not something you do by hand.
So the phrasing hypothesis is parked, unproven, and I am telling you about it because the version where I do not tell you is how most case studies get written.
How I count the engines that report nothing
You cannot get a report out of ChatGPT, Claude or Perplexity, and you may not have one for Google either. You can still count, and counting properly is most of the job. This is the protocol I run, written down before the data arrives.
Twelve money queries, six engines, fresh incognito on every one. ChatGPT, Perplexity, Gemini, Google AI Mode, Claude, and Copilot. Copilot is on that list for a reason that took me an embarrassing amount of time to see: it is the only engine where I can check my hand-count against a real export. Leaving it out meant my manual method had no calibration anywhere in it. Now every baseline reconciles the Copilot column against the Clarity and Bing numbers for the same window, and if those two disagree, the problem is my counting, not the engine.
Every cell scored 0, 1 or 2, with a competitor column beside it, first response only. Not "did I feel visible," but a number I can re-run.
That incognito rule is not decoration. A logged-in session carries your history, your location, and on several engines a memory of what you asked before. Ask about your own industry from an account that has spent months researching your own business and you will get a flattering answer no customer will ever see.
Two limits on it I would rather state than have pointed out. Incognito clears cookies, not IP geolocation, and for a local business location is the dominant personalisation axis, so the answer is still weighted to where you are sitting. And some engines require an account, so logged-out is not achievable everywhere. Treat a logged-out run as a control for comparing your own runs over time, never as a simulation of your customer, because your customer is signed in.
Then two passes almost nobody runs. The first is a memory-versus-search test: ask the same prompt with web search turned off. If a business still surfaces, it lives in the model rather than being retrieved, and those are different problems with different fixes. The second is an echo audit. Before I report that a business is named in eight sources, I check whether those eight are independent or one press release repeated eight times. Eight echoes of one origin is just a one.
And the read is pre-committed. Before I look at anything, the success threshold for the ninety-day re-run is written down and signed off. What counts as signal, what counts as null, what counts as inconclusive. The instrument is frozen, so no query gets swapped for a friendlier one later. A metric chosen after seeing the data is p-hacking with extra steps.
I expect most of those engines to come back at or near zero for my mom's shop. I will publish that. The whole reason to pre-commit a threshold is so a disappointing number still gets reported. Otherwise it becomes a different metric on the way to the page.
What this means for your business
Three asks, in order of how much money they save you.
Ask what instrument produced your number. Not just what tool, what instrument. "Our platform tracks AI visibility" is not an answer. "Bing Webmaster Tools AI Performance, exported monthly" is an answer.
Ask how many runs sit behind it. If somebody sampled an engine once per query, their margin of error is larger than almost any effect they are claiming about your site. Three runs per query per engine is the floor I use, and it is barely enough.
Ask what the number is a share of. Share of authority sounds enormous until you learn it only counts the queries your business already appears on. Mine is about a third, and that third is measured on the handful of questions the shop already wins, not on the market.
What I sell, and what I am not showing you
I take on a small number of audits, so I benefit if you find this credible. Read your way through it knowing that.
And by my own standard I owe you the export. The twelve queries, the dated runs, the raw counts, and the pre-committed thresholds are going up on your side of the table alongside the ninety-day results rather than being described at you. If that page does not exist by the time you read this, the honest reading is that I asked you to hold me to something I have not met yet.
Questions readers ask me about measuring AI visibility
Which AI engines can I actually measure for free?
Three tiers, checked 28 July 2026. Tier one is Microsoft Copilot. Bing Webmaster Tools and Microsoft Clarity both report your citations, free, with the grounding queries attached. Tier two is Google. Search Console began showing impressions from its AI surfaces in June 2026, but the rollout is staged and many sites do not have it yet. Tier three is ChatGPT, Claude and Perplexity. Those three report nothing to site owners at all. Any number you are shown for them came from sampling, so ask for the run count.
Can I track ChatGPT citations for my website?
No, because OpenAI publishes no citation report for site owners. Server logs get you closer than most people realise. OpenAI runs separate agents for separate jobs: GPTBot trains, OAI-SearchBot handles search appearance, and ChatGPT-User makes the live fetch during an answer. A ChatGPT-User hit is a warmer signal than a training crawl. What no log tells you is whether that page made it into the answer. Past the logs, the honest method is sampling. Ask ChatGPT your customer questions on a schedule, log every run, and report the spread.
Does Google Search Console show AI Overview citations?
Not citations, no. In June 2026 Google began rolling out generative AI performance reporting in Search Console. What it shows is impressions from Google's AI surfaces, not a list of the answers that quoted you. The rollout is staged, so check your own property before planning around it. AI Overview impressions and clicks have also long been folded into your ordinary Search totals without being broken out.
Which AI visibility tools actually report first-party data?
For citations, two. Bing Webmaster Tools has an AI Performance report, and Microsoft Clarity has an AI citations dashboard. Both cover Microsoft Copilot and its partner surfaces, and both are free. For impressions, Search Console's generative reporting, where your property has it. Everything else covering ChatGPT, Perplexity or Claude is running its own sampling and showing you the result. That can be worth paying for when the run count is visible. It is not first-party data and should not be priced as though it were.
How many times should I test a query before trusting the answer?
Three runs per query per engine is my floor, and it is barely enough. On my own run set, one query returned 0% and then 69% on consecutive attempts, and per-cell standard deviations ran as high as 49 points. A single run tells you what happened once. It does not tell you what usually happens, and the gap between those two is where most AI visibility claims live.
Does running queries in incognito show me what a customer sees?
No, and this is worth being precise about. Incognito drops your account history and your past questions. That is what stops an engine flattering you with a personalised answer. It does not drop your IP location, so a local business still gets a locally weighted result. Your real customers are also signed in. Treat a logged-out run as a control for comparing your own runs over time, never as a picture of what a customer sees.
Does blocking Google-Extended keep me out of AI Overviews?
No, and it is one of the most common mix-ups in GEO advice. Google-Extended governs Gemini apps and Vertex AI grounding. Google's own documentation states it does not affect inclusion in Google Search. AI Overviews are part of Search, so the levers there are the ordinary snippet directives: nosnippet, max-snippet and data-nosnippet. Blocking Google-Extended is still a real choice, because it does drop you out of Gemini and Vertex grounding.
What does share of authority actually mean?
It is the share of cited sources that belong to you, counted only on the queries where you already appear. If an answer cites six sources and two are yours, your share on that query is about a third. It says nothing about how many conversations you are absent from. Read it as depth on the questions you already win, never as coverage of your market.
Run the same check on your own store
The free measurement stack, including the GA4 setup most people miss, is in the 30-minute measurement stack guide. To find out whether the engines can read you at all before you worry about citations, run the 10-minute readability check. If none of these terms are familiar yet, start with the plain-English GEO guide. A side-by-side of what each engine does and shows you is in the engine comparison.
The full case study behind the citation numbers, with its limits stated, is here. And if your problem is that your best wins never show up in analytics at all, that one is here. You can check your own robots.txt against the free scanner.
Run the free AI Readiness Scanner: paste in your robots.txt and your homepage source, get a scored report on whether ChatGPT, Perplexity, and Google AI can actually read your business. About two minutes, no signup, and nothing you paste ever leaves your browser.
Run the free scan → Get The Receipts free ▸
I run the SEO and AI visibility for my family's Shopify shop and publish the receipts, good and bad. Every number on this site comes from a named tool export, and corrections get published rather than edited away.