Our methods, our samples — including the results that flatter nobody.
Everything LiamVi recommends traces to something we measured, and everything we measured is described here: how the test worked, how big the sample was, and what it cannot prove. These are measurements, not laws. The dataset is small and early — we say so in every article — and it grows with every weekly sweep. When a conclusion changes, we change it here, in the open.
Last updated 12 August 2026.
How we measure
Citation harvests
For a set of real keywords, we ask an assistant the query and record which pages it actually cites — then compare those against the pages that rank on Google for the same query but were ignored. Cited versus ignored, same keyword, is the cleanest comparison we know how to run.
Pre-registered framing tests
Before running the test we wrote down the prediction — which query framings should trigger source retrieval, and in what order. Pre-registration keeps us from quietly picking the result we liked after the fact.
Matched pairs
To test whether domain authority decides citations, we compare pairs of pages competing on the SAME keyword — one from a higher-authority domain, one from a lower — and count who actually got cited. Same query, same moment, different authority: only the variable under test differs.
Blind judging
Our own scoring meters are calibrated against independent reviewers who judge article quality cold — one reviewer per article, no scores shown, no comparison — so a meter only survives if it agrees with a judgement it could not have copied. Stated plainly: those cold reads were done by independent reviewer agents on a seven-dimension scorecard, not yet by a person. A human blind read of the same set is prepared and unfinished; we will publish it when it is done, whatever it says.
The current dataset
Stated plainly: this is an early dataset — enough to change our own product decisions, not enough to call laws. Our public goal is 1,000 keywords harvested. We keep this page updated as the dataset expands, and we will flag the milestone here when we cross it.
| Keywords harvested | 33 | cited vs ignored pages, mostly ChatGPT — goal: 1,000 |
| Framing ladder | 5 topics × 4 framings | pre-registered before the test ran |
| Matched authority pairs | 21 | same keyword, different domain authority |
| Articles scored | 241 | 43 cited vs 198 ignored, 800+ words |
| Blind-judged calibration | 15 articles | judged cold, no scores shown — 9 carried our score, 6 a competitor's |
| Published articles benchmarked | 44 | where real professional content scores on our meters |
What we found so far
Question framing decides whether retrieval happens at all
In the pre-registered ladder, none of the conceptual “what is X” framings triggered source retrieval on the engine we measured, while every dated, comparison-shaped framing did. Limit: five topics, one engine primarily; engines differ — Perplexity, for example, cites far more freely. The direction was unambiguous; the exact rates are not portable.
Domain authority did not decide citations
Across 21 matched pairs, the higher-authority domain won 11 and lost 10 — statistically nothing. Our weekly sweeps keep finding small, recently created sites holding citation slots beside household names. Limit: 21 pairs cannot rule out a small effect; they rule out a decisive one in this sample.
The unit of citation is the passage, not the page
Some assistant citations link to a single highlighted sentence inside a page rather than the page itself — one article in our harvest was cited four times for four different passages. This is why our editor measures sentences an AI could lift, not pages.
Repeating what already ranks is not what gets cited
What we measured. Across the 33-keyword harvest we recorded 276 ranking pages and 73 pages ChatGPT actually cited for those same queries. Only 9 pages appeared in both sets — 3.3% — and just 12.3% of citations came from Google’s top ten. Seven citations in eight came from outside the first page of results. The two funnels barely overlap, so a page assembled from the consensus of everything already ranking is competing in the wrong set.
What the engine says about itself. Asked to describe its own source selection, ChatGPT told us it often ignores pages that add no new evidence — five articles repeating one press release are not five independent sources. That is a self-description, not a measurement: models confabulate plausible mechanisms, and we would not publish it on its own. We publish it because it matches what we measured first, and because it named the defect in our own product — our early scorer rewarded matching the top-ten consensus, which is the set our harvest shows is nearly disjoint from what gets cited.
Limits. 33 keywords, three related industries, mostly ChatGPT. The overlap counts are pages, not domains, and one harvest is one moment — these are the odds we observed, not a rule about the web.
Harvest and engine self-report measured 29 July 2026 · published here 12 August 2026.
The null we published against ourselves
We believed answering the questions searchers ask (Google's People-Also-Ask) would earn citations. It did not: across the 33-keyword harvest, cited pages answered 58.8% of searcher questions and ignored pages answered 58.2% — a difference of nothing. We kept the measure in our editor because it remains our strongest signal of article QUALITY against the blind judges’ ratings (15 articles) — but we stopped claiming it buys citations, and our product copy says so.
Our own content score pointed the wrong way
The finding. In a pre-registered head-to-head trial, 15 articles were each read cold by an independent reviewer who saw no scores at all. Our fitted content score ran negativeagainst those judgements: rho −0.61 across the 9 articles our score covered. The competitor score we benchmarked did the same thing on its 6 (rho −0.63), and combined across all 15 the figure is again −0.61. Inside a keyword pair, the scores named the judge’s winner once in six tries: one draft scored 90 and drew a judge’s 81, while a 78 drew an 86. This is a category-wide flaw in scores fitted to what ranks — ours included, and ours first.
What the two sample sizes mean. Both numbers in our records are real and count different things: 15 is the number of articles blind-judged; 9 is how many of them carried our score (the other 6 carried the competitor’s, because each article was written under one stack or the other). So rho −0.61 (n=9) is our meter against the judges, and rho −0.61 (n=15) is the whole trial, each article measured against whichever score it carried.
Scope — read this part. Every draft in the trial scored between 71 and 90. What failed was not the meter’s ability to tell bad writing from good; it was the meter’s ability to rank good writing against itself. Small n, one narrow band. A confound we logged rather than buried: our editing loop revised each draft until the score cleared 75, so the weaker-scoring drafts received more editorial work than the strong ones. The cold reads were done by independent reviewer agents, one per article, on a seven-dimension scorecard; a second format — a three-lens side-by-side panel over 9 pairs — agreed with the first on 7 of those 9. A human blind read of the same articles is prepared and not yet done.
Why it is on this page. It is the reason we stopped selling a score as a quality grade and rebuilt ours as a coverage checklist — what you are missing, not how good you are. Any tool that fits a number to the consensus of the ranking pages should be asked for this measurement. We ran it on ourselves and published the answer.
Measured 28 July 2026 · published here 12 August 2026.
Where real winners score
The 43 pages ChatGPT actually cited score a median of 65 on our citability checklist — and only 5% reach 85. Real published professional articles score a median of ~76 on our SEO-terms meter (44 measured). This is why our score bands call 55–84 “strong” and treat 85+ as informational, never a target: chasing points past the winners’ zone produces the stuffing assistants discard.
What this cannot prove
- Causation. We observe what cited pages share; we cannot yet prove that adding those properties causes citation. The weekly citation tracker is accumulating exactly the outcome data that will test this.
- Other engines.Most measurements are ChatGPT; Gemini, Perplexity and Google’s AI surfaces are sampled far more thinly and behave differently.
- Other markets. SERP data (volumes, questions, term usage) is measured on US-English Google as a global-English proxy.
- Any promise. No score, ours included, is a probability of being cited. Anyone selling you a guaranteed citation is selling you the dashboard, not the outcome.
The commitment: the sweeps run weekly and we keep this page updated as the dataset expands — the numbers above change, the last-updated date changes with them, and when a conclusion stops surviving the data we correct it here, in the open, not quietly. At 1,000 harvested keywords we re-run every test on this page and publish what held.
And it grows with you. Every keyword a customer tracks and every draft they score adds measured ground truth to the same dataset — which pages win citations, which properties those pages shared, what changed a week later. The more you write and measure through LiamVi, the sharper the insights and the advice we can hand back — to you, and to everyone measuring beside you.