LiamVi

← All articles  ·  10 September 2026

An upright measuring instrument with two channels of the same size on its face. The left is an outlined slot holding a full run of navy graduation ticks; the right is a solid orange bar with no outline and no graduations. Beneath, two words stand one above the other: Observed, then Inferred.

AI tools with the best generative engine optimization features: what the features do, and what they cannot tell you

We expected that answering the questions searchers ask would get a page quoted by an AI assistant. It seemed obvious, and it is the assumption underneath a lot of generative engine optimization tooling currently being sold. Across our 33-keyword harvest, cited pages answered 58.8% of those questions and ignored pages answered 58.2%. That is a coin flip. We published it anyway, because it was our own assumption that broke, and the method, the sample and the limits are on our research page.

The useful part is not that a single metric failed. It is why it failed. Question coverage is a property of your page, and any tool that can read HTML can measure question coverage to two decimal places. Whether an assistant quotes you is decided somewhere else — in a step that happens before your page is under consideration at all. A product can be perfectly accurate about the first thing and know nothing about the second.

So when you are reading a comparison table of generative engine optimization features, the question that separates them is not which list is longer. It is what each feature can actually see. This article sorts GEO capabilities into three tiers by that test, shows the measurements behind the sort, quotes what Search Central publishes about its own AI surfaces, and describes five AI visibility platforms in their own words. There is no ranking here, and the last section explains why we think ranking these tools would be dishonest rather than merely difficult.

What is generative engine optimization, and how is it different from SEO?

Generative engine optimization is the work of getting your page mentioned or quoted inside an AI-generated answer rather than ranked in a list of links. The short form is GEO. The adjacent term is AEO — answer engine optimization — and the two are used almost interchangeably across the vendors selling either.

The largest search engine with published documentation does not accept the distinction. Search Central's guidance on optimizing for generative AI search states the position plainly: "From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO." The guide names both acronyms first — "'AEO' stands for 'answer engine optimization' and 'GEO' for 'generative engine optimization'" — and then declines to treat either as a separate discipline. That scope is narrow and it matters: the page describes AI Overviews and AI Mode, and it carries no authority over ChatGPT, Claude or Perplexity.

Our own boundary sits elsewhere, and it is about measurement rather than craft. SEO and GEO may well be the same writing job; they are not the same measurement job. A rank is a position in a list, checkable by anyone who runs the query. A citation is a link an assistant shows inside an answer to support what it just said, and the decision to show it is invisible from your side of the wire. That is how the two vendors with published documentation describe it: Google's systems review "the specific information from those retrieved pages" and then show "prominent, clickable links to relevant web pages that support the information in the response", and OpenAI's help pages describe a citation as a source you select to open. Neither describes copying a sentence out of your page, so we do not claim they do. What we have measured is narrower and more useful: some citations we collected point at a single highlighted sentence inside a page rather than at the page itself, and one article in our harvest was cited four times for four separate passages. Whatever happens inside the model, the thing being pointed at is often a passage. A page-level score is not so much the wrong number as the wrong unit.

That is why the GEO label is worth arguing about briefly and no longer. What you are buying when you buy an AI visibility platform is not a new discipline. It is a set of instruments, and instruments differ in what they can see.

What does a generative engine optimization feature actually do?

Every visibility measurement feature in the five platforms we opened on 1 September 2026 does one of three jobs, or some combination of them. A GEO tool reads your page, it tracks what an assistant said about your brand when asked a question, or it watches which AI crawlers reached your server. Most of these platforms do two or three of those jobs under a single dashboard, which is why the feature lists read longer and more varied than the underlying measurement work is.

Two kinds of feature sit outside those three jobs, and a reader using the sort needs to know where it stops. A prompt-demand dataset measures other people's questions rather than your visibility: Profound's "Prompt Volumes" offers to "See what millions of people ask AI, and align strategy with demand", and the question to ask about a feature like that is about the vendor's sample of conversations, not about your site. Drafting and workflow features are outside the sort too — Writesonic's platform names an AI article writer and an action centre beside its tracking. Both kinds may be useful and we have measured neither; neither is an instrument pointed at your visibility, and a comparison table that lists them in the same column as one is counting different things.

The vendors' own descriptions make the three jobs easy to spot. Otterly.AI sells "AI Search Analytics" to "Monitor how your brand & website appears (or disappears) across every major AI search engine" — monitoring the answer surface — and separately a "Content Audit" promising "Crawlability checks, Content Audit & Prediction, and Content Briefs", which is the page-reading job. Scrunch splits its own menu the same way, into "Monitoring & Insights" and "Agent Experience", where "Site Diagnostics" offers to "Improve how AI consumes your site" and "Agent Traffic" offers to "Analyze AI website traffic". Profound lists "Answer Engine Insights" to "See how AI represents your brand in every conversation" beside "Agent Analytics" to "Track how your site is interpreted and crawled by ChatGPT, Gemini, Claude, Perplexity, and more". Every quotation here is from the vendor's own page, opened on 1 September 2026, and this wording changes often enough that you should open the pages yourself before buying.

None of that makes the three jobs equivalent. The difference between them is not quality. It is what is knowable.

What can a GEO feature observe, and what can it only infer?

Two of the three jobs are direct observation: reading your page, and reading your own server logs. Monitoring the answer surface is a sample, and the retrieval decision itself can only be inferred. A dashboard renders all three in the same typeface, which is why the sort below is worth running on any GEO tool's pricing page before you buy — it takes about five minutes.

Three identical plates stand apart on open ground, each named above it. Tier one holds two outlined blocks marked Your page and Your server; Tier two three identical blue chips beside Answer surface; Tier three is empty, with a small orange block to its right marked Retrieval decision.
Tier one reads things you already own — your page and your own server logs. Tier two samples what an assistant said at one moment. Tier three is a decision made off your site: none of it happens on your page, so no product reads it there. The difference between the three is not quality; it is what is knowable.

Tier one: directly observed — your page and your server

A tier-one feature reads something you already own and reports what it found. That means your page's word count and heading structure, the schema markup on it, and which AI crawlers requested which URLs and what status code each one received. Nothing in that list is estimated, so there is nothing to trust: the tool is reading something that exists and telling you what it says. A first-tier number that is wrong is a bug, and you can catch the bug by opening the page yourself.

The crawler half of this tier is the most under-used check a practitioner can run, because the vendors publish exactly what each robot is for. OpenAI's crawler documentation separates two user agents by job. OAI-SearchBot "is used to surface websites in search results in ChatGPT's search features", while GPTBot "is used to crawl content that may be used in training our generative AI foundation models." The consequence is stated in the same table: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." The two controls are independent — "a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot" — and the documentation notes that for search results, "it can take ~24 hours from a site's robots.txt update for our systems to adjust."

That distinction is worth checking tonight, and the check is log analysis you can run without buying a tool. A team that blocked the training crawler in 2024 to opt out of AI may have left the search crawler untouched, or may have blocked both without meaning to, and either way the evidence is sitting in a file that team already owns. This is OpenAI's documentation about OpenAI's crawlers; other assistants publish their own, and we are not extrapolating from it.

Tier two: sampled — the answer surface

A tier-two feature asks an assistant a question at a moment and records the answer. What it records is which brands were mentioned, in what order, and which sources the assistant cited. Brand mention tracking is the job most of these platforms lead with, and monitoring of that kind is genuinely useful: it tells you when your brand stopped being mentioned for a question your buyers ask.

It is also a sample. The same question tomorrow, from another account in another country, can return a new answer, and an AI visibility number built from a hundred tracked prompts is a number built from a hundred tracked prompts.

A second limit on tier two is documented rather than inferred. Search Central's guide describes "query fan-out" as "a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query", with its own example: a user asking "how to fix a lawn that's full of weeds" may generate fan-out queries such as "best herbicides for lawns" and "remove weeds without chemicals". The prompt you typed into a brand tracker may not be the query that actually ran. SE Ranking's SE Visible advertises a "Query fan-out analysis" to "See the sub-queries behind each AI answer" and marks that analysis "coming soon" on its own site as of 1 September 2026, which tells you both that the gap is recognised and that it was not closed there on the day we looked.

One solid orange bar high in the frame, marked You typed, with three thin navy lines running down from it into three navy-outlined bars of clearly unequal length below, marked What ran.
Google documents query fan-out — a set of related queries its model generates to answer yours — and OpenAI says its search typically rewrites your query into one or more targeted queries. So the prompt a brand tracker types may not be the query that actually ran. How many are generated is not fixed, and we have not measured how much of our own ranking-overlap gap it explains.

Tier three: inferred only — the retrieval decision

A tier-three number is an inference about a decision made off your site. The decisions in question are whether a search ran at all, which passage of your page an assistant points at, and why a competitor's brand was mentioned and yours was not. No product reads any of that off your page, because none of it happens on your page.

Start with whether a search runs. OpenAI's current help documentation says "ChatGPT may search the web automatically when your question would benefit from current information" — a decision made from the question, not from your site. We ran a pre-registered test on exactly that: five topics, four framings each, prediction written down before the run. None of the conceptual "what is X" framings triggered source retrieval on the engine we measured, and every dated, comparison-shaped framing did. The ladder was pre-registered and is published with its limits, and those limits are narrow: five topics, one engine primarily, and engines vary — Perplexity cites far more freely than the engine we tested. The direction was unambiguous; the exact rates are not portable to your market.

The same boundary is documented in Search Central's guidance. Grounding works "by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index", and eligibility for an AI Overview or an AI Mode answer has two conditions, neither of which lives in your copy: "a page must be indexed and eligible to be shown in Google Search with a snippet", and "a site must be included in Search generative AI features in Search Console". The guide then adds the sentence most product pages leave out: "Just because a page meets all requirements, best practices, and complies with the policies, doesn't mean that Google will crawl, index, or serve its content. Indexing and serving aren't guaranteed."

The sort matters because tier three is where most of the promise in this category lives and where none of the observation does. A tier-one check tells you what your page says. A tier-two tracker tells you what an assistant said on Tuesday. Neither can tell you why, and any number that claims to is a model of the decision rather than a reading of it.

Why can a tool score your page perfectly and still be wrong about citations?

A score can be exactly right about your page and still tell you nothing about whether an assistant quotes it. Citation is not decided at the page level, and we know that from the inside: we built a score that got the direction backwards, and we published the result.

The coin flip is the cleanest version. We believed that answering the questions searchers ask — the People Also Ask set on a search results page — would earn citations. Across our 33-keyword harvest, cited pages answered 58.8% of those questions and ignored pages answered 58.2%, a difference of nothing. We kept the measure in our editor because it remains our strongest available signal of article quality against the blind ratings we hold, and we stopped claiming it buys citations. That harvest covers 33 keywords across three related industries, mostly ChatGPT. It is small and early, our sweeps run weekly, and we publish what changes; the harvest and its limits are on the same research page.

The second result is about our own instrument. In a pre-registered head-to-head trial, 15 articles were each read cold by an independent reviewer that saw no scores at all, and our fitted content score ran negative against those judgements: rho −0.61 across the 9 articles our score covered, and −0.61 again across all 15 in the trial. Rho is a rank correlation, and a negative one here means the plain thing it sounds like: within that set, the drafts our score liked best were the ones the cold reader liked least. The benchmarked competitor meter did the same thing on its own 6, at −0.63 — a category-wide flaw in scores fitted to what ranks, ours included and ours first, which is why we are not naming the other product. We publish the trial's method and both sample counts. Three limits ride with that finding and we publish all three. Every draft in the trial scored between 71 and 90, so what failed was the meter's ability to rank good writing against itself rather than to tell bad writing from good. The cold reads were done by independent reviewer agents on a seven-dimension scorecard, not by people, and a human blind read of the same articles is prepared and not yet done. And our editing loop revised each draft until the score cleared 75, so the weaker-scoring drafts received more editorial work than the strong ones — an honest alternative explanation for the direction we found. We rebuilt our scoring around that boundary instead of treating the result as a law, and the full account of what a score can and cannot grade is in our piece on when a tool's score lies.

A third measurement makes the direction plausible rather than a fluke. Across the same 33-keyword harvest we recorded 276 ranking pages and 73 pages ChatGPT actually cited for those queries. Nine pages appeared in both sets — 3.3% — and 12.3% of citations came from Google's top ten. Seven citations in eight came from outside the first page of results. Those counts sit on our research page with their limits: they are pages rather than domains, and a single harvest is a single moment. A third limit belongs here and it is the one an attentive reader reaches for first. Our ranking set was harvested for the query we typed, and an assistant does not necessarily search the query you typed: OpenAI's help pages say its search "typically rewrites your query into one or more targeted queries that it sends those providers", and Google documents query fan-out on its own AI surfaces. So some of that gap may be pages ranking for a fan-out query we never harvested rather than pages the ranking consensus missed, and we have not measured how much. Our read, and we mark it as a read rather than a measurement: the overlap is small, and small in the same direction as the other two results, so a meter fitted to the consensus of what already ranks is not measuring the set that gets cited, and a high reading on that meter is not evidence of the thing a buyer cares about.

Our read, and we mark it as a read rather than a measurement: a GEO score that outputs a single page-level number is doing something genuinely useful — finding omissions — while presenting the number as something it is not, a probability of being quoted. The fix is not to distrust the score. It is to know which tier it came from.

What does Google say about optimizing for generative AI features?

Search Central's published guidance names four things sold as GEO capabilities that do nothing on Google's AI surfaces. The page is the guide to optimizing your website for generative AI features on Google Search, last updated 10 July 2026, and its scope is stated in its title: AI Overviews and AI Mode. Search Central is not describing the other assistants, and neither are we when we quote it.

The document carries a section headed "Mythbusting generative AI search: what you don't need to do". Four items in that section map onto features currently being sold:

  • Machine-readable files for AI. "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." The guidance is even-handed about it: an `llms.txt` "will neither harm nor help your site's visibility or rankings in Google Search", and "It's completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files."
  • Chunking. "There's no requirement to break your content into tiny pieces for AI to better understand it." The same passage adds that "There's no ideal page length", which is worth holding beside any tool that recommends a target length.
  • Rewriting for machines. "You don't need to write in a specific way just for generative AI search. AI systems can understand synonyms and general meanings of what someone is seeking".
  • Structured data. "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add." Structured data is still recommended — "it's a good idea to continue using it as part of your overall SEO strategy, as it helps with being eligible for rich results" — but as ordinary SEO, not as an AI capability.

The same page addresses the tools directly, and both halves of what it says matter: "Be wary of third-party tools that promise ranking success or claim to use 'internal' Google metrics. No third-party tool has access to our internal ranking or AI systems. Use these tools if they help your workflow, but be sure to evaluate their advice against our official guidance." That is not advice to stop using measurement platforms. It is a statement about where the ceiling sits, from the only party in a position to state it.

For its own surfaces the guidance names a first-party alternative: "To measure how your content is performing in generative AI features on Google Search and Discover, use the Generative AI performance report in Search Console." That report is free and it is first-party, and what it reports is impressions — "how many times links to your site were shown to a user in a generative AI feature on Google Search", broken out by page, country, device and date. The report does not hand you the answer text, the prompt behind it, or where your brand sat in the list, and it covers Google's own AI features and nothing else — which is precisely why third-party brand monitoring exists for the rest.

A gap we will not paper over: that documentation speaks for a single search engine. We looked for equivalent current statements from OpenAI, Anthropic and Perplexity about machine-readable files and chunking on 1 September 2026 and did not find them, and we have not measured the question ourselves. So the honest scope of everything in this section is one search engine's AI surfaces, and anyone telling you what `llms.txt` does inside another assistant is telling you something neither that vendor nor we can currently show.

Which AI tools have generative engine optimization features today?

SE Visible, Profound, Otterly.AI, Scrunch and Writesonic are the five products we opened and read on 1 September 2026. Each is described below in its own words, and each sells some combination of the three jobs above. This is our scope, not a census of the market, and it is not a ranking. We hold no measured, like-for-like comparison of these platforms' quality, and we know of nobody who has published one.

  • SE Visible, from SE Ranking, offers to "Find out where your brand appears, how it compares, and what shapes its visibility across ChatGPT, Gemini, Perplexity, AI Mode, and AI Overviews." Its features menu lists query fan-out analysis, local AI visibility and brand sentiment tracking, each marked "coming soon" on the day we looked.
  • Profound describes itself as "the full stack marketing platform for the marketer of the future" and names eight assistant surfaces across the top of its page, from Perplexity and Claude to Google AI Overviews. Its "Prompt Volumes" offers to "See what millions of people ask AI, and align strategy with demand", and "Agent Analytics" to "Track how your site is interpreted and crawled" by those assistants' crawlers. A free report is advertised.
  • Otterly.AI calls itself "the Content Intelligence Platform for AI Search" and offers to "Analyze how AI search engines mention, and cite your brand - across ChatGPT, Google AI Overviews, Claude, Perplexity, Google AI Mode, Gemini, and Copilot." Its "Content Audit" is unusually direct about the question it is answering: "Know why AI skips your content — and fix it."
  • Scrunch sells the crawler side hardest, splitting the platform into "Agent Experience" and "Monitoring & Insights". "Content Delivery" offers to "Deliver an optimized experience to AI agents", "Site Diagnostics" to "Improve how AI consumes your site", and "Monitoring & Citations" to "Know how your brand shows up in AI."
  • Writesonic leads with the workflow rather than the dashboard: "See where AI ignores you. Fix it with content, citations, and technical work. Prove the lift in traffic, pipeline, and revenue."

And the disclosure, since this piece is a judgement about the category we sell into. We build LiamVi, which tracks which sources hold the citations for your keywords, scores drafts on the properties cited pages share, and flags keywords where measured retrieval is zero. We are small and new. By this piece's own test our checklist is a tier-one instrument with a tier-two input, and you should hold it to the questions in the next section exactly as you hold the others. What we can hand you is our research — the methods, samples and limits behind everything we publish, including the result two sections up that flatters nobody.

This category turns over quickly. Products rename, merge and add capabilities between our weekly sweeps, and the vendors selling GEO as a done-for-you service are a separate set again — we described the companies who sell it as a service elsewhere. Treat any list, ours included, as a dated snapshot to check on the vendor's own page.

How do you compare generative engine optimization features without a ranking?

The honest comparison is not a count of checkmarks but a question about tiers. Sort each capability into its tier, then ask the question that tier can answer. Read what follows as our view of what matters rather than a neutral standard: we build a product in this category, and criteria written by a vendor are always shaped by what that vendor built.

For a tier-one check, ask to see the raw evidence, not the score. If a tool says your page is missing something, that tool should show you the page state it read — the headings it found, the schema it parsed, the crawler requests it logged, with timestamps. A first-tier feature that will only show you a number is withholding the easiest thing in the product to show.

For tier-two brand monitoring, ask how often the platform samples, across which engines and regions, and what the run-to-run variance is. This is the question with the most purchase, because two honest trackers sampling at separate moments will disagree, and their disagreement is not evidence that either is lying. Ask us the same thing: we have not published a variance figure for our own sampling, and until we do, that is a limit in our product rather than a rhetorical device.

For anything that reaches into tier three, ask what the number infers and from what. If a platform reports a likelihood, a readiness score or an AI visibility score for being quoted, the useful follow-up is what that score was fitted to. If the answer is "the pages that currently rank", you now know exactly what our own rho result was about, and you can weigh the insight accordingly.

Two limits on the sort itself, stated where the sort is. It tells you what a feature can know; it cannot tell you whether the vendor implemented that feature competently, and those are separate questions. And it is a framework rather than a measurement: we have not scored these tools against each other, we do not intend to publish a ranking we cannot support, and if anyone does publish a like-for-like accuracy analysis of AI visibility platforms, that document will be more useful than this one.

What people actually ask us

These four questions arrive most often from readers comparing AI visibility tools. We answer two of them and refuse two, and the reason travels with each.

What is the best AI for search engine optimization? We will not name one, and the reason is not diplomacy. We hold no like-for-like measurement of these platforms against each other, and as far as we can find neither does anybody else, so any single name would be a preference dressed as a finding. There is also no best generative engine optimization feature in the abstract: the best instrument for a business with a technical SEO backlog is not the best instrument for a business that already ranks and is not being mentioned. What we can give you is the test — sort the capabilities by tier, ask the three questions above, and buy the tool whose answers are specific.

Can ChatGPT do SEO? A language model can help with two jobs that are not the same, and it is much better at one of them. Drafting, outlining and rewriting are things a model does; whether the result gets quoted is not something the model can tell you, because the retrieval decision sits outside the page it just wrote. Our own evidence for that gap is the coin flip above: a page property we could measure precisely separated cited from ignored pages by 0.6 percentage points across 33 keywords. For measurement, match the instrument to the question: Search Console will tell you whether links to your pages were shown inside Google's AI features, and a brand tracker will sample what an assistant actually said about you. Neither of them will tell you why.

What are the top five AI tools right now, and what are the top five SEO tools? We are refusing both questions rather than answering them badly, and here is the whole reason: a top-five list implies a comparison was run, and we did not run one. The five vendors described above are those whose pages we opened on 1 September 2026, not the best five of anything. If you need a ranking today, look for a ranking whose author publishes the method and the sample — and if you cannot find that, treat the absence as information about the category rather than about your search.

Are there free generative engine optimization tools? Yes, with limits worth knowing. The Generative AI performance report in Search Console is free, first-party, and reports on AI Overviews and AI Mode only. Several of the platforms above advertise free trials or free reports on their own pages: Profound advertises a free report, Otterly.AI and SE Visible advertise free trials, and Scrunch advertises a free AI visibility audit. Those are their published claims as of 1 September 2026, not something we have tested. What we have not found free anywhere is the thing tier three would need — a documented, sampled analysis of why one page was retrieved and another was not.

Sources