
How to get cited by ChatGPT in 2026: eight signals, and what has actually been tested
The page that currently ranks first for this question sells eight signals, scores them out of sixteen, and tells you what your total means. We went through those eight asking a narrower question than whether they work: has anybody measured them? Four of the eight have been tested by somebody other than us, each with a sample and a comparison group — and three of those four came back against the advice the page gives. The page cites none of that work.
That is the part worth sitting with. The evidence behind these signals is not missing. It is missing from the checklist. A page that scores eight items out of sixteen and attaches two thresholds to the total published neither the tests that support its items nor the tests that contradict them, and the difference between a measured claim and a confident one is visible once you go looking.
One thing to clear away first, because this question has two very different readers. If you are here because you want to cite ChatGPT as a source in an essay or a paper, that is a separate subject with its own formats and rules, and we have written it up in how to get citations from ChatGPT. Everything below is about the other question: getting your own page quoted when someone else asks ChatGPT something.
What the playbook claims, and what it does not show
The top-ranked page for this question names eight signals and scores them out of sixteen. It attaches two thresholds to that score and publishes no sample behind either one. The page is a GEO playbook dated 26 April 2026 and last updated on 14 June 2026, and on the rubric it says: "We score each on a 0-1-2 rubric and use the totals to prioritize where the budget goes." Eight signals at two points each makes sixteen, and the page attaches two numbers to that total — "Most sites we look at score under 6 out of 16" and "anything over 12 starts showing up in citations within a quarter."
It is worth saying plainly that this is a serious page. It is specific where most writing on this subject is vague, it is more detailed than the rest of the results we read for this question, and at least one of its sub-rubrics is genuinely operational: it tells you to score each H2 section on "complete-on-its-own (0/1), under 150 words (0/1), specific claim with numbers or names (0/1)." That is a thing a person can actually do on a Tuesday.
What the page does not contain is a sample. Reading it for a method, a comparison group, or a count of what was examined returns nothing on any of the three. It says "The patterns are visible if you look at thousands of citations across queries", which describes looking rather than reporting what the looking found, and it points at its own statistics page; it cites no external study and no third-party dataset. The one quantified line about its own work — "Most sites we look at score under 6 out of 16" — is a statement about scores the page assigned, not about citations it counted, and it carries no n.
We are not saying its author measured nothing. We cannot know that. We are saying that what would let a reader check the second number — anything over 12 starts showing up in citations within a quarter — is not on the page, and that a scoring threshold is exactly the kind of claim that needs one.
Which of the eight signals has anyone actually measured?
Five of the eight have some published measurement behind them, two have none our search reached, and one is not a claim about citation at all. Here is the whole list against the published evidence we found as of 10 September 2026.
| Signal (its name, from the playbook) | What it claims | What has been measured | |---|---|---| | Source domain authority | the single biggest predictor of citation | measured — and it did not hold. Our 21 matched pairs: higher-authority page won 11, lost 10. And 82,108 citations in which the top authority tier converted at 15.0%, below every other tier | | Chunk-level citability | the engine pulls a chunk, not a page | passages are demonstrably selected. Citation links carrying the quoted span; one article cited four times for four passages. Not a test of what makes a passage selectable | | Schema density | markup is the cheapest win available | measured twice, null both times. 1,006 pages against a control population (OR 0.678, p = .296); 1,885 pages that added JSON-LD against roughly 4,000 controls | | Entity clarity | you must be a recognisable entity to the graph | no published test we found | | Freshness signals | engines penalise stale content | measured — and it runs their way. 16.975 million cited URLs average 1,064 days old against 1,432 for organic results | | Expert presence | named authors get cited more than anonymous ones | measured — no across-the-board lift. Bylines added to 123 pages against untreated controls | | Citation reciprocity | a trusted source citing you raises confidence in you | no published test we found | | Active monitoring | you cannot optimise what you cannot measure | not a claim about citation at all — see below |
Read the table as our reading of somebody else's checklist. We build a tool in this category, we decided what counts here as a published test — a sample and a comparison group — and two of its rows are graded against our own work. What the third column is not is a census of everything published on the subject: it is what looking turned up.
Source domain authority: the signal we measured that did not hold
Higher domain authority won 11 of 21 matched pairs and lost 10. The pairs were pages competing on the same keyword, one from a higher-authority domain and one from a lower, and the signal the playbook ranks first is the one that measurement contradicts. It is also the strongest claim the page makes for any of the eight: "The single biggest predictor of whether an AI engine cites you is whether the engine has decided your domain is a credible source." In the one place we could check it, the biggest predictor did not separate the winners from the losers at all.

The limits belong in the same breath as the result, because they are real. The pairs were matched on the keyword and on nothing else, so two pages on two domains also differ in their content, their links, their age and how squarely they answer the question; the 11–10 is a count of what happened rather than a controlled test of authority. Twenty-one pairs cannot rule out a small effect; they rule out a decisive one in that sample. The full write-up, with the sample and the limits, is on our research page and in more detail in what we tested and what came back null.
Our 21 pairs are not the weight of the evidence here, and a far larger dataset points the same way. The highest authority tier, DA 80 to 100, turned retrieval into a citation 15.0% of the time — below every other tier, which ran between 21.5% and 23.6% — across 548,534 pages ChatGPT retrieved and the 82,108 citations that came out of them, in a study AirOps published on 12 March 2026. Their sentence is "High-authority sites were retrieved often but cited at a lower rate than every other tier." On a sample ours cannot reach, the signal the page ranks first is not doing the work there either.
Our weekly sweeps keep turning up recently created sites holding citation slots beside household names. That is an observation across sweeps rather than a second test, and we record it as one.
Chunk-level citability: the signal our evidence supports
Some AI citations write the quoted sentence directly into the link. That makes passage-level selection the one item on the playbook's list our own evidence supports, and the evidence is physical rather than inferred. Those citations point at a single sentence inside a page rather than at the page. One article in our harvest was cited four separate times, each citation pointing at a different passage in it. One page, four selections.
The mechanism is a published web standard, not a theory of ours. The `#:~:text=` form is specified in URL Fragment Text Directives, a Draft Community Group Report dated 13 December 2023, whose stated purpose is that "text directives add support for specifying a text snippet in the URL fragment" so that "the user agent can quickly emphasise and/or bring it to the user's attention." When a citation arrives in that form, the span is written into the URL in plain sight.
Be precise about what that proves. It shows a system selected a span of words. It does not tell us why that span was chosen over another one, it does not mean every assistant works this way, and it does not mean the page around the passage counts for nothing — a page still has to be retrieved before a passage in it can be picked.
Here the playbook and our measurement agree, and we would rather say so than manufacture a disagreement. Its definition is close to what we would write: "A chunk is typically a paragraph or a small group of paragraphs under a single heading. The engine pulls that chunk into its response, attributes it to your URL…" Its three-criteria rubric — self-contained, short, carrying a specific claim with numbers or names — is a reasonable description of the spans we actually saw quoted. Of the eight, this is the item we would spend on first — a judgement rather than a result, held because the evidence that passages get selected is the most direct evidence on the list, not because anyone has tested whether the rubric changes what gets selected.
The other five: three have been tested, two have not
Schema density, freshness signals and expert presence each have a published test behind them; entity clarity and citation reciprocity have none our search reached. Not one of those tests appears on the page that ranks first, in either direction.

Schema density — tested twice, and null both times. The claim: "Schema.org markup is the cheapest GEO win available." Schema presence came back null — an odds ratio of 0.678 at p = .296 — across 730 AI citations from ChatGPT and Gemini, 75 commercial queries and 1,006 unique pages, with Google's top-ten organic results as the control population, in a February 2026 preprint by Kurt Fischman. The correction is the interesting part: the raw association is, in the paper's words, "a methodological artifact: Google's ranking algorithm systematically enriches top-10 organic results for schema-bearing pages." It does report one exception, and it is the paper's finding rather than our recommendation: pages carrying Product or Review schema with concrete attributes like prices and ratings were cited at 61.7% against 41.6%, p = .012. Then somebody ran the intervention. 1,885 pages added JSON-LD between August 2025 and March 2026 and their citations barely moved — −4.6% in AI Overviews, +2.4% in AI Mode, +2.2% in ChatGPT, measured against roughly 4,000 control pages thirty days either side of the change, in an Ahrefs study published on 11 May 2026. They say they cannot confidently attribute the −4.6% to schema itself, and they call the two positives indistinguishable from zero. Their sentence is "Adding schema produced no major uplift in citations on any platform." Both of those tests come from companies that build tools in this category, and where a competitor's work holds up we would rather say so than leave it out.
Freshness signals — tested, and this one runs the page's way. The claim: "AI engines penalize stale content. The penalty is sharper than in traditional SEO because the engines are explicitly trying not to surface out-of-date answers." URLs cited by AI assistants run 1,064 days old on average, against 1,432 days for the organic results on the same queries — 16.975 million cited URLs across seven platforms, in an Ahrefs study published on 28 July 2025. Their sentence: "The average age of URLs cited by AI assistants is 1064 days, compared to 1432 days for URLs in organic SERPs—25.7% 'fresher'." That is the largest comparison on this list, and it points where the page says it points. Be exact about what it is, though. It compares the ages of two populations of URLs; it does not hold a page constant and vary its stated dates, so it cannot tell you what happens when you update yours, and it is not a measurement of a penalty. It is also worth reading the absolute number rather than the gap: 1,064 days is close to three years. Cited content is fresher than organic results. It is not fresh.
Expert presence — tested, and honestly inconclusive. The claim: "Anonymous content gets cited less than named-author content. Named-author content gets cited more when the author has a verifiable public profile." Detailed author bylines — headshots, credentials, expertise claims and schema — added to 123 blog pages produced no across-the-board citation lift against untreated pages on the same blog, and a −1.3% effect on AI retrieval crawl once the sitewide trend was accounted for, which is statistically zero: that is Seer Interactive's byline test, in their words "The bylines didn't deliver the across-the-board citation lift we expected, but some of the signals showed positive indicators." Bing was the one place something moved: citations to the treated pages roughly doubled period over period against a gain of about a third on the control group — though the write-up gives that change and not the citation counts underneath it, which is the question this article is asking of everyone. The authors call the whole thing "not conclusive." One thing to note about that write-up, in an article about checking: it carries no publication date, and an undated source is a weaker source.
Entity clarity. The claim: "To get cited as an authority on a topic, you have to be a recognizable entity that the graph associates with that topic." Untested by us, and we found no published test. It is also the hardest of the eight to test honestly, because "recognisable to the graph" has no agreed measurement — the test would have to define its own instrument before it could run, and then defend the instrument.
Citation reciprocity. The claim: "If a source the engine already trusts cites you, the engine's confidence in you goes up." Untested by us, and we found no published test of it. Link counts have been measured against citation counts, which is a different object: how many domains link to you is not whether a source the engine already trusts citing you changes anything. It is also the most expensive of the eight to test, because the intervention takes months and the comparison group has to sit still while it happens.
One honest word about those last two rows. "No published test we found" is a report of what looking turned up on 10 September 2026, not a fact about the world. We cannot prove a negative about what has been published, this is one desk's search, and if you know of a test we missed we would rather be told than leave the gap standing. What we can say without qualification is what the page itself does not do: it cites none of this evidence, in either direction, while attaching a score to every item on its list.
What the checklist leaves out entirely
Every one of the playbook's eight signals is a property of a page — its domain, its chunks, its markup, its author, its links, its dates — and the decision that moved our own citation rate furthest is not a property of a page at all.
In a pre-registered test we held five topics constant and phrased each of them four ways: conceptual, entity-state, dated, and a recommendation with a jurisdiction attached. None of the five conceptual framings drew a citation from anyone. Every dated framing did. That decision — how the question is framed — is made when the topic is commissioned, before a page exists to have properties, which is why no page-level checklist can reach it. The sample and its limits are on our research page, and the fuller account is in the strategies write-up.
This is not the playbook's freshness signal, and the two must not be run together. Its freshness signals item is about markup on a page stating when the page was published or updated, and the study above measures how old cited pages are. Our result is about how the underlying question is phrased. Those are different objects, and our finding says nothing whatever about theirs — treating our ladder as evidence for or against their freshness item would be inventing a result, which is the failure this whole article is about.
The gap runs the other way too. Three of the four strategies we did test are not on the eight-item list at all.
Question coverage. Across a harvest of 33 keywords, the pages ChatGPT cited and the pages it ignored answered searcher questions at effectively the same rate. That was our own hypothesis before it was our own null, and we published it against ourselves.
Ranking position. 9 of the 73 pages ChatGPT cited also sat in Google's top ten for the query that produced the citation. That figure is the overlap between two sets rather than a citation rate by rank: we never measured how often a page outside the top ten was cited against how often a top-ten page was, so nothing in it says ranking well makes a citation less likely.
Crawler permissions. Whether a domain carried an AI-crawler block did not separate the cited domains from the ignored ones in our harvest. The permission that actually governs the question is narrower: OAI-SearchBot is the crawler OpenAI documents as surfacing sites in ChatGPT's search features, and it was blocked by 1 of the 230 domains we looked at — too little variation to test at all.
So the checklist and the published evidence barely overlap in either direction. The page cites none of the tests that exist for its own items, and three of the four strategies we tested ourselves are not on its list. That is the honest state of this subject in 2026, and it is worth knowing before you spend a quarter against a scoring rubric.
Active monitoring is not a signal, it is how you find out
Measuring your own citations makes no claim about what causes one. That is why the eighth item on the playbook's list is also the only one that needs no test. "If you cannot measure your citations, you cannot optimize them," the page says, and we agree without reservation — it is simply advice to go and look.
Worth knowing what looking involves. Measuring citations means asking assistants real questions on a schedule and recording which pages come back in the answer, which is a different measurement from an impressions report: a search console can tell you a page was shown somewhere, but it cannot tell you whether an AI answer quoted you. In our own harvest, asking the same assistant the same query twice could return another answer and another set of citations, so any honest number here comes with a sampling method attached. We wrote up what a visibility measurement actually counts, and where the numbers come from, in what an AI visibility dashboard counts.
How to read any citation checklist, including ours
The fastest way to weigh an AI-citation checklist, ours included, is to ask what the denominator is behind each item on it. That single question separates a measured claim from a confident one faster than reading the rest of the page, and it works on every vendor in this category.

Three questions do most of the work. What was measured? Not "we looked at thousands of citations", but what was counted and on which queries. Compared against what? A claim that cited pages have a property is empty until you know that uncited pages have it less often. How many? A fraction, not a percentage — "11 of 21 pairs" tells you what "52%" hides.
You should apply that to us. We build a tool in this category, so read our reading of these eight as our view of what matters rather than as a neutral standard. Our own scoring instrument has published problems of its own: tested against blind quality reviews run by independent reviewer agents, the drafts our meter rated highest were the ones the reviews rated lowest, at rho −0.61 across the nine articles it covered. When a conclusion changes we change it in public.
And an untested item is not a wrong item. Entity clarity may well matter. Three months of outreach may well be the best thing on the list. What changes once you know which items are untested is not whether you do them, but how much certainty you buy with your budget — and that trade-off belongs to you, with your own site and your own quarter. We cannot make it for you, and a scoring rubric that adds an untested item to a tested one and returns a single number out of sixteen is quietly making it for you.
What we cannot tell you
This article is a map of the published evidence behind one checklist, not a verdict on what works. Six of its eight signals are untested by us; four of those have been tested by other people, and two have no test we could find at all. If somebody publishes a real test of entity clarity next month, it belongs in the measured column of that table and we will put it there.
Our own numbers are small and early. The harvest behind them is 33 keywords across three related industries, mostly ChatGPT, with a much thinner sample on other assistants; engines differ, and Perplexity cites far more freely than the one we measured most. Twenty-one matched pairs is twenty-one, and the harvest grows with every weekly sweep — we keep measuring, and we publish what changes. These systems also move: the retrieval behaviour we measured is a policy as much as a property, and thresholds can change without notice, which means every number above is dated rather than settled.
What you can do this week costs nothing. Take the queries you actually want to win, ask an assistant each of them with search on, and read who sits in the sources. Then take your own page on that subject and read it asking one question — which sentence here could be quoted whole, on its own, and still say something specific? That check needs no rubric and no budget, and it is the item on the list whose evidence you can see most directly — the quoted span written into the link, where anyone can read it.
Sources
- Winston Digital Marketing — How to get cited by ChatGPT in 2026 — the eight signals, the 0-1-2 rubric and the score thresholds quoted above. Published 26 April 2026, updated 14 June 2026; verified 10 September 2026.
- Kurt Fischman — Does schema markup predict AI citation? — 730 AI citations, 75 queries, 1,006 pages, Google's top ten as the control population; schema presence null at OR 0.678, p = .296. A February 2026 preprint — the author's own research page dates it 12 February 2026 and the aiXiv posting 22 February 2026; verified 10 September 2026.
- Ahrefs — We tracked 1,885 pages adding schema — the JSON-LD intervention against roughly 4,000 control pages, and the three platform figures. Published 11 May 2026; verified 10 September 2026.
- AirOps — The influence of retrieval, fan-out and Google SERPs on ChatGPT citations — 548,534 retrieved pages, 82,108 citations, and citation conversion by domain-authority tier. Published 12 March 2026; verified 10 September 2026.
- Ahrefs — Do AI assistants prefer to cite fresh content? — 16.975 million cited URLs across seven platforms, against the organic results for the same queries. Published 28 July 2025; verified 10 September 2026.
- Seer Interactive — Do author bylines influence AI visibility? — bylines added to 123 pages against untreated controls, and the authors' own "not conclusive". The page shows no publication date; verified 10 September 2026.
- W3C Community Group — URL Fragment Text Directives — the `#:~:text=` syntax and its stated purpose. Draft Community Group Report, 13 December 2023; verified 10 September 2026.
- LiamVi — Research: our methods, samples and limits — the matched pairs, the passage evidence, the framing ladder, the question-coverage null and the ranking overlap, each with its sample. Last updated 12 August 2026; verified 10 September 2026.
- LiamVi — Generative engine optimization strategies for AI visibility — the four tested strategies and their published limits. Verified 10 September 2026.
- LiamVi — How to get citations from ChatGPT — citing ChatGPT as a source, in academic formats. Verified 10 September 2026.
- LiamVi — Best AI visibility analytics for search optimization — what a citation measurement actually counts, and the run-to-run variance behind any figure. Verified 10 September 2026.
- OpenAI — Bots and crawlers — GPTBot and OAI-SearchBot, and what each is used for. Verified 10 September 2026.
- LiamVi — Content optimization tools in 2026 — our own scoring instrument tested against blind reviewer-agent judgements, and what it got wrong. Verified 10 September 2026.
- LiamVi — AI and search engine optimization in 2026 — impressions are not citations, and why we measure citations directly. Published 12 September 2026.