
AI visibility analytics for search optimization: what the numbers count, and what a score is measured against
The 43 pages ChatGPT actually cited in our keyword harvest reach a median of 65 on LiamVi's citability checklist, and only about one in twenty of those 43 reaches 85. We put that figure on our research page because it is uncomfortable for us: a very high reading on our own meter is not what a cited page looks like. It is also the most useful thing we know about AI visibility analytics, because it tells you what kind of object a number in these platforms actually is.
An AI visibility analytics platform hands you two main kinds of number, and each fails in its own way.
The first kind is a tracking number — brand mentions, citations, share of voice. It is a sample of what a set of chosen prompts returned from chosen AI engines on chosen days. The second kind is a score — a content, optimization or readiness rating. It is a model fitted to a target somebody picked, and the useful thing to ask is which target. A third kind, sentiment, is neither, and we come to it below.
This guide is about reading all of them. It ranks no tools and names no winner; we build a platform in this category ourselves, which is precisely why we are not going to hand you a league table. Read the rest knowing that, and hold us to the same standard we are asking you to hold everyone else to.
What is an AI visibility analytics number actually counting?
Every headline figure in an AI visibility analytics platform is a tally of assistant answers, not of web pages: what a chosen list of prompts returned from chosen AI engines, on the days the tool ran them.
The three tallies almost every tool shows are brand mentions, citations and share of voice. We have written up what an AI visibility tool measures and how to judge them separately, so the definitions live there and this guide does not repeat them. What those definitions leave out is the part you need for reading a figure: each tally is a count over a prompt list, and that prompt list is an editorial decision somebody made. Usually you made some of it during setup, in about ten minutes, and never looked at it again.
What an AI search monitoring product actually runs
An AI search monitoring product runs three things on your behalf: a list of prompts, a schedule, and a set of engines. Those three choices produce the number. The prompt list decides which queries get asked; the schedule decides how often, and therefore how much of the engines' own movement you see; the engine set decides whose behaviour is being measured at all, and the engines do not behave alike.

A monitoring tool that sampled ChatGPT daily and Gemini weekly would not be measuring one AI visibility figure — it would be combining two readings taken at unequal densities. The cadence belongs beside the number the way a sample size belongs beside a survey result, and some vendors do publish it — usually in their documentation, not on the feature list you are comparing. Three products' own pages, read on 4 September 2026, each say daily: Profound's knowledge base says its prompts are "the queries Profound automatically sends to answer engines on a daily basis"; Otterly.AI's help centre says it "monitors your brand visibility across all major AI search engines on a daily basis"; Peec AI's documentation says "we run your prompts across AI platforms like ChatGPT, Gemini, and Copilot daily". That is three products read on one date, not a survey of the category, and it is the check, not the conclusion: go and find the cadence in the documentation before you read the number. None of this is hidden malice; it is what scheduled sampling of a moving system looks like. But a tracking product that will not tell you its cadence is asking you to read a figure with half the label torn off.
Change the prompt list and the number changes, with no change at all to your website. Add ten queries your buyers ask in a market where your brand is strong and your share of voice rises. Add ten from a market you have never written for and it falls. Nothing about your content moved. So "our AI visibility went up 12 points" is a claim about your content only if the keyword list stayed still, and most teams cannot say that theirs did.
Which surfaces does the number cover, and is Google AI Overviews among them?
An AI visibility number covers only the answer surfaces the platform samples, and the coverage list belongs beside the figure as much as the sample size does. Across the product pages we read for our own data-accuracy guide in August 2026, the surfaces named most often were ChatGPT, Gemini, Claude, Perplexity and Google AI Overviews — and the lists were not the same from one product to the next. That is a handful of vendor pages read on two dates in August 2026, not a survey of the category.
Google AI Overviews are the odd one in that list, because they reach someone who never opened an assistant at all — an answer served above traditional results to a person doing an ordinary query. We think that makes the coverage question a question about your buyers rather than about the technology: a platform that samples Claude and Perplexity closely and Google AI Overviews thinly is a good fit for some markets and a poor one for others. We have not measured which markets, and we are not going to guess.
Here is our own limit on this, stated where the claim is: our measurements are mostly ChatGPT, with Gemini and Perplexity sampled thinly and Google AI Overviews and Claude barely at all. So we can tell you that surface coverage varies by product and belongs on your checklist; we cannot tell you how Google AI Overviews or Claude select sources, because we have not measured them. Anyone who does tell you should be asked for the sample.
There is a second gap, and it sits inside the word "citation"
In our own harvest, some assistant citations pointed at a highlighted sentence inside a page instead of at the page — one article was cited four times for four separate passages. We wrote that finding up in how to get citations from ChatGPT, and it matters here because a page-level count and a citation are not the same object. A cell reading "3" may mean three pages, or a page quoted three times, and a platform that does not distinguish them is asking you to guess.
Why does the prompt list decide the number?
In our own pre-registered test, the framing of a tracked prompt decided whether the AI engine consulted the web at all — and an engine that never consults the web cites nobody, which arrives in your analytics as a zero.

We tested this on ourselves before we believed it, and wrote the prediction down before the run: topic held constant, only the framing varied. We recorded the result on our research page in these words: "none of the conceptual 'what is X' framings triggered source retrieval on the engine we measured, while every dated, comparison-shaped framing did." The published limit travels with that finding wherever we use it: five topics, one engine primarily, and the engines diverge — Perplexity, for example, cites far more freely. The direction was unambiguous. The exact rates are not portable, and we do not offer them as a law about the web.
The engines say something similar in their own documentation, which is where you should verify it instead of taking our word. Google's Grounding with Google Search documentation, last updated 2026-09-03, states that a grounded response carries metadata with source links, and then: "However, there are several reasons this metadata might not be provided, and the prompt response won't be grounded." OpenAI's October 2024 post introducing ChatGPT search puts it from the other side: "ChatGPT will choose to search the web based on what you ask, or you can manually choose to search by clicking the web search icon." The decision sits inside the engine and is made before any page is considered.
A tracked keyword list built mostly out of "what is" prompts can therefore hand you a flat zero that says nothing whatsoever about your content. That zero is a fact about the prompt rather than about the page. And if your analytics platform cannot separate "nobody was cited" from "no search ran", two unlike situations arrive in the same empty cell — and only the first is a problem you can write your way out of.
Practically, this makes your tracked prompt list an instrument and not a settings page. Open it and read it that way. How many of those prompts are conceptual? How many name an entity, a comparison, a year, a market? You are not gaming anything by asking; you are finding out which of your queries can produce a measurement at all. It is also the cheapest improvement available in most accounts, because it costs no content work and changes what the tracking can see.
Why do two honest tools report different numbers for the same brand?
Two AI visibility tools can both be measuring honestly and still disagree about the same brand, because in our own harvest asking the same assistant the same query twice could return another answer and another set of citations.
We published that finding in our own guide to data accuracy in this category, and it changed how we read a disagreement between two read-outs. The scope goes wherever the claim goes: a 33-keyword citation harvest collected mostly from ChatGPT, with Gemini and Perplexity sampled thinly — a small, early sample. Two platforms tracking the same brand at unlike moments, with their own prompts and their own run counts, will not land on the same figure even when both are measuring well. The variance is a property of the thing being measured rather than a defect in the measurer.
This has a direct edge for anyone comparing tools during a trial. Running two platforms side by side for a week and treating the gap between them as an accuracy test tells you very little, because you have not held the sample constant. What the gap does tell you is how far apart the two tools sample, which is worth knowing and is separate from which is right.
Whether a particular vendor's data is right is its own test with its own method, and that same guide sets it out, including the only cross-check we have run ourselves, on Surfer's citation data, which came back positive. This guide is about the figure in front of you.
What is the score inside the dashboard measured against?
LiamVi's citability checklist is built from the properties cited pages share, and the 43 pages ChatGPT actually cited in our harvest reach a median of 65 on it.
We measured it and published the whole finding, with its samples attached, on our research page: those 43 cited pages reach a median of 65 on the citability checklist, and only about one in twenty of them reaches 85. Real published professional articles reach a median of about 76 on our SEO-terms meter, across 44 measured. Our bands call 55–84 strong and treat 85 and above as informational instead of a target, and the figures above are why we drew them there. The step from that distribution to "chasing points past the winners' zone produces padding" is our reading of it rather than a measured effect: we have not tested whether a higher score costs a page a citation. We observe what cited pages share; we cannot yet prove that adding those properties causes one.
Read these as our view of what matters, shaped by the product we built rather than as a neutral standard. That is the honest framing for any criteria list written by somebody selling in the category, ours included.
The general shape survives being lifted out of our data. A content score is a floor and not a finish line. If the pages that won the citations sat in the middle of a meter, a page in the high eighties on that meter is not thereby a better page — it is a page outside the range where the cited pages in our sample sat. A team optimising toward 95 is optimising toward a region our evidence does not describe, which is not the same as a region we have shown to be worse.
That is not a claim that these meters are useless. It is a claim about which instrument you are holding. A coverage checklist tells you what you are missing, and missing things is worth knowing. It does not tell you how good the writing is, and it is not a probability of being cited. We say that last part plainly about our own product: no score, ours included, is a probability of being cited, and anyone selling you a guaranteed citation is selling you the dashboard, not the outcome.
We know the difference between those instruments because we got it wrong first. Our early meter rewarded matching the vocabulary and coverage of the pages already ranking — the standard content optimization approach across traditional SEO tools, and a reasonable-looking idea until you test it. We then measured that meter against independent cold reads of the same articles, and it did not agree with them. We rebuilt the thing as a coverage checklist — what you are missing, not how good you are — and stopped selling it as a quality grade. The full write-up, including the part that flatters nobody, is on our research page.
A detail about that calibration, because it is the kind of thing a vendor usually smooths over: those cold reads were done by independent reviewer agents on a seven-dimension scorecard, not yet by a person. A human blind read of the same set is prepared and unfinished, and we will publish it when it is done, whatever it says.
The reason this belongs in a guide about analytics and not in a guide about writing is that the meter usually sits on the same screen as the tracking numbers, in the same typeface, looking like the same kind of fact. It is not. The first is a sample of what engines did; the second is a model fitted to a target somebody chose, and the target is almost never printed beside the figure. Ask what the meter was fitted to. It is a fair thing to ask and it has an answer.
And the third kind of number: sentiment
Sentiment is a model's judgement of tone, which makes it neither a count nor a fitted meter, and it deserves to be read as the judgement it is. Some platforms report a sentiment reading beside the mention count — our data-accuracy guide describes a product in this category doing exactly that, alongside visibility and position.
The honest framing is that a sentiment figure is a second model's opinion of a first model's prose. That is not worthless: if an assistant describes your product as expensive across forty answers, you would want to know. But it inherits every problem above — the prompt list, the sampling cadence, the surface coverage — and adds one of its own, because tone is a judgement call and reasonable readers disagree about it. We have not measured sentiment accuracy in any platform, ours or anyone's, so we have no number to offer you here and will not pretend otherwise. Treat a sentiment trend as a prompt to go and read the answers, not as a performance metric in its own right.
The same caution applies to any blended "performance" figure. A performance score that folds mentions, citations, position and sentiment into a single headline is a weighted opinion about what matters, and the weights are the vendor's, not yours. Ask to see them, and ask what happens to the total when one input is missing.
What can no visibility number see?
No AI visibility number can see an answer where no engine ever looked, and a page-level count cannot see which sentence inside a page an assistant actually quoted.
The first blind spot is covered above. The second explains a mistake that costs money: the assumption that a traditional SEO ranking report can stand in for AI visibility analytics.
The two funnels barely overlap in our data. Across the same 33-keyword harvest we recorded 276 ranking pages and 73 pages ChatGPT actually cited for those queries, and 9 pages appeared in both sets — 9 of the 276 that ranked, and 9 of the 73 that were cited. Seven citations in eight came from outside the first page of Google's results. The limits sit on the page with the finding and we carry them here: 33 keywords, three related industries, mostly ChatGPT; the overlap counts are pages and not domains; and the harvest is a snapshot of a moment. These are the odds we observed, not a rule about the web.
What that means for your weekly read-out is direct. Position data and citation data were near-disjoint sets in our sample, so rank tracking cannot stand in for AI visibility tracking. A strong traditional SEO report is not evidence that assistants are quoting you, and a platform that blends both into a single headline figure hides which of the two moved. It is also why a generative engine optimization (GEO) tool samples answers instead of deriving them from ranks: the ranks were not where the citations were.
The third blind spot is arithmetic: share of voice hides its denominator
Share of voice is a ratio, and every ratio has a denominator you are not being shown. It compares your brand against a competitor set across a list of tracked prompts — so it moves when a competitor's mentions move, when the keyword list changes, and when the answers that cited nobody are handled another way.
That last case deserves naming, because it is invisible and it is large. When a tracked prompt returns an answer citing no sources at all, the slot did not exist that day. A monitoring platform can treat that cell in two defensible ways: drop it from the average as no-data, or count it as a competitor loss. Both are honest choices, and they produce opposite trend lines from identical raw observations. Two tools can therefore show unlike share-of-voice figures for the same brand without either being wrong, purely on that decision. Ask which your platform makes, and ask what proportion of your tracked queries fell into it last month — a share of voice computed mostly over empty slots is competitor benchmarking in name only.
Which questions should you ask about the number?
Five things are worth asking about an AI visibility figure itself: its prompt list, its run count, its variance, what any score was fitted to, and what a zero in the cell means. They work on any AI search analytics product, and they interrogate the number rather than the vendor; the vendor's own honesty and the cleanliness of its input data are a different test, and we published that one on its own. Read the list as ours, shaped by the platform we build, and not a neutral standard. We have answered all five for ourselves below, including where our answer is thin.
1. What prompt list produced this figure, and who wrote it? You need the list, not a count of it. Ours is the set of keywords a customer chooses to track, which means the instrument is partly built by the person reading it — worth knowing on the day the number moves.
2. How many runs, over how long? A figure from a single day's run is a snapshot of a system that does not repeat, which is the argument for continuous AI visibility monitoring over a one-off audit. Our own sweeps run weekly, and the dataset grows with them.
3. How much does it move between runs? This is the variance point and the likeliest to be met with silence. A read-out whose numbers never move should worry you more than one that wobbles. We have observed run-to-run variance in citations on identical queries; we have not published a variance figure of our own, and until we do, that is a gap in our answer rather than a fact about the category.
4. What is the score fitted to — and where do the pages that actually got cited land on it? This separates a calibrated meter from a confident one. Our answer is above: a median of 65 on the citability checklist, across the 43 cited pages in our harvest, with about one in twenty of them reaching 85. We have not surveyed the category, so we will not claim that nobody else publishes an equivalent. What we will say is that a vendor who cannot answer it has not tested their meter against the outcome you are buying it for.
5. What does a zero in this cell mean? Nobody was cited, or the engine never ran a query — separate cases that call for opposite responses, and a platform that cannot tell you which has handed you a cell you cannot act on. The feature that matters most here is whether the tool separates those two at all.
A figure that survives all five is not necessarily right, but it is checkable, and checkable is the entire difference between a measurement and a claim. Ask them before you sign, and ask them of us.
What insight does a tracking number actually support?
An AI visibility figure supports fewer conclusions than its precision implies, and knowing which ones is the difference between an insight and a decoration. Three hold up in our experience of our own data, and they are worth more than the decimal places.
The first is direction over time on a fixed keyword list — if the list did not change and the monitoring cadence did not change, movement across several weeks is a real signal, and a single week's move usually is not. The second is which prompts are empty, because an empty slot is an editorial finding about the query and it points at work that content cannot do. The third is which competitors keep appearing where you do not, and the insight you can act on there is not the count itself: it is what those cited pages contain, which you get by opening them.
What a tracking figure does not support is a precise cause. It cannot tell you that last month's article earned this month's brand mentions, because the engines moved, the prompt list may have moved, and the variance sits underneath both. A GEO platform that offers you causal insight from sampled counts is offering an interpretation rather than a measurement, and that gap matters most exactly when the number is going your way. Most of the insights a report advertises are the honest kind; the causal ones are where to slow down, and no feature label on the screen will tell you which insight you are looking at.
How do you check your own dashboard this week?
Putting five of your tracked queries to two AI engines, twice each, across two separate days gives you 40 observations. That is four per prompt-and-engine pair — a spot-check rather than a variance study, and still more than most dashboards will show you about their own movement.
That arithmetic — 5 prompts × 2 engines × 2 runs × 2 days — costs nothing but about an hour of attention spread across a week, and it needs no paid tools, because the assistants themselves are free to ask. Pick the two engines your buyers actually use; if that is ChatGPT and Claude, use those, and if Google AI Overviews matter more to your market, run the query in ordinary Google as your second surface. Use the prompts exactly as your platform has them, phrased the way a buyer asks, and record three things per run: whether the engine cited any sources at all, whether your brand appeared, and which competitors did.
One limit belongs here, not in a closing caveat: you are running these in a consumer app, and that is not the same instrument as a platform's scheduled run. OpenAI's help page on searching the web with ChatGPT says that "if memory is enabled, ChatGPT may use relevant saved memories when rewriting a search query", that ChatGPT "may use an approximate location based on your IP address to provide relevant local results", and that in a managed workspace you should "ask your administrator whether web search is enabled for your role". Your account travels with your answers, which makes what you collect a reading of your own instrument, not a replacement for the platform's.
Then read what you collected against the tool's own view of the same period, in this order:
- Did any prompt return no sources on either day? Those are queries where the slot did not exist when you looked. Verify how the platform handled them. That is point five, answered from your own runs.
- Did the same prompt return other sources on day two? That is movement you saw yourself, on your own queries, from four observations of each — not a variance figure, and four runs cannot establish one. Write down how large it was. What it is good for is that it stops you reading a small weekly change as a result.
- Does the tool's direction match what you saw? Not the exact figure — the direction. If your own reading and the platform's disagree about which way the week went, that is worth a support conversation before it is worth a switch.
The common and quite honest outcome is that the two broadly agree, the variance is visible, and nothing is wrong. That is a good result: you now know roughly how much your number moves for reasons that have nothing to do with your content, which is the context every later reading needs. The uncomfortable outcome — a report that never moves at all while your own runs disagree with each other — is the case worth asking about.
What people ask us
Which AI visibility tool is the best one? We will not crown one, and we would not believe anybody who did without publishing their method. Which tool fits depends on the AI surfaces your buyers actually use, on whether you still need traditional rank tracking beside it, and on how much of the sampling method the vendor will show you. We wrote up what these tools measure and how to judge them so that the judging stays yours.
Which AI should a team optimise for — or write with? This question carries two meanings with separate answers. If you mean which assistant to write with, we have not measured that and will not guess. If you mean which engine to optimise for, our measurements are mostly ChatGPT with Gemini and Perplexity sampled thinly, and the honest answer is that the engines behave in their own ways enough that the choice is a matter of where your buyers are. The single thing our data does say is that the framing of a prompt decides whether any engine looks at the web at all.
Is LLM optimization a different thing from GEO? Generative engine optimization (GEO), answer engine optimization and LLM optimization are three labels for the same observable work: which queries trigger retrieval, which pages get quoted for them, and which sentences inside those pages get lifted. Nobody outside the labs can observe a model's internals, so anyone selling optimisation of the model itself is selling something they cannot measure. Optimising what a model can retrieve is real work; optimising the model is not a thing you can buy, and a GEO feature list that implies otherwise is describing an ambition rather than a measurement.
How is GEO measurement different from traditional SEO reporting? GEO measures your presence in an answer; traditional SEO measures where you sat on a results page. In our sample those were nearly disjoint sets of pages, which is why GEO analytics sample answers instead of deriving them from ranks. The practical distinction for a buyer is that traditional SEO has an obvious ground truth you can look at, while AI search visibility has a sampled one — so the sampling method is the feature you are actually buying, and the first thing to ask a GEO vendor about.
Is AI visibility analytics one product, or a stack? For most teams it is a stack, and the honest reason is that no product we know of covers every AI answer surface your buyers use at the density you would want, and in our own measurements Google AI Overviews and Claude are the two we cover most thinly. The best answer available to you is the one you can verify: pick for the surfaces that matter, keep traditional rank tracking if rankings still earn your traffic, and treat any blended performance figure as the summary it is.
Is SEO dead now with AI? No, and our own numbers are why we can say it calmly: in the 33-keyword harvest the ranking pages and the cited pages were nearly disjoint sets — 9 pages in both, out of 276 that ranked and 73 that were cited. Two funnels, not one replacing the other. Traditional rankings still send the traffic they send, and the work that earns them has not stopped working. What changed is that a second funnel now exists, it is measured another way, and a report about the first tells you very little about the second.
What we cannot tell you
Our citation dataset is 33 keywords across three related industries, measured mostly on ChatGPT — small and early, and every claim above inherits that limit. The rest of the arithmetic, so you can weigh it yourself: 43 cited pages against 198 ignored ones among 241 articles scored, 21 matched pairs, and SERP data measured on US-English Google as a global-English proxy. It is enough to change our own product decisions and not enough to call laws.
We have not shown causation. We observe what cited pages share; we cannot yet prove that adding those properties causes a citation, and the weekly tracking is accumulating exactly the outcome data that would test it. We have not measured Gemini, Perplexity, Claude or Google AI Overviews anywhere near as deeply as ChatGPT, and they behave in their own ways. We have not measured sentiment accuracy at all. We have not published a variance figure of our own, which is the gap in our answer to point three above.
And no score, ours included, is a probability of being cited.
The sweeps run weekly, the dataset grows, and when a conclusion stops surviving the data we correct it in public rather than quietly. If the median moves, we will say so — including if it moves somewhere less convenient for us than 65.
Sources
- LiamVi — Research: our methods, our samples — the 33-keyword citation harvest, the pre-registered framing ladder, the citability checklist medians, the ranking-versus-citation overlap counts and the published limits. Page last updated 12 August 2026; verified 4 September 2026.
- LiamVi — The best accurate data platform for AI search optimization in 2026 — run-to-run variance, sampling honesty, the sentiment-and-position reporting described there, and the Surfer cross-check. Published 31 August 2026; verified 4 September 2026.
- LiamVi — The best AI visibility tools in 2026: what they measure, and how to judge one — brand mentions, citations and share of voice defined, and how to judge them. Verified 4 September 2026.
- LiamVi — How to get citations from ChatGPT — that article's finding, in its words: the unit of citation is the passage, not the page. Verified 4 September 2026.
- Google Cloud — Grounding with Google Search — grounding metadata may not be provided and the prompt response may not be grounded. Page last updated 2026-09-03; verified 4 September 2026.
- OpenAI — Introducing ChatGPT search — the assistant chooses whether to consult the web based on what you ask. Post dated 31 October 2024, with in-page updates through 5 February 2025; checked 4 September 2026.
- Profound — Answer Engine Insights overview (knowledge base) — the vendor's own statement of its sampling cadence: prompts are the queries it sends to answer engines on a daily basis. Verified 4 September 2026.
- Otterly.AI — How often does Otterly.AI check AI search engines? (help centre) — daily monitoring across the engines in the account. Page dated 5 August 2026; verified 4 September 2026.
- Peec AI — Welcome to Peec AI (documentation) — prompts run daily across the AI platforms named there. Verified 4 September 2026.
- OpenAI — Searching the web with ChatGPT (help centre) — saved memories may be used when rewriting a search query, an approximate location from the IP address may be used, and in a managed workspace web search may not be enabled for a role. Verified 4 September 2026.