How to Get Cited by AI Search Engines: ChatGPT, Perplexity, and Google AI Overviews
Last updated: August 31, 2026 · By Joseph Olivas, Founder, MEAN Consultors · 9 min read
I get asked about this in almost every SEO conversation now, usually phrased as a worry: our traffic is going to AI answers, so how do we at least be the source they quote? It is a fair question, and it has a better answer than most of what is circulating online. Google published official guidance on this in 2026, and a striking amount of it consists of telling site owners to stop doing things.
How a page actually becomes a citation
Start with the mechanism, because most bad advice comes from misunderstanding it. Google describes two techniques behind its generative AI features. The first is retrieval-augmented generation (RAG), also called grounding, which uses core Search ranking to retrieve relevant pages from the index and then generates an answer from what those pages say. The second is query fan-out: the model takes one question and generates a set of concurrent related sub-queries to widen what it retrieves.

Figure 1: The retrieval path Google documents for its generative AI features. Citation happens at the end of a chain that begins with ordinary indexing.
Two consequences fall out of this immediately. First, being cited requires being retrieved, and being retrieved requires being indexed and eligible to show with a snippet in ordinary Search. There is no separate AI index you can submit to. Second, because of fan-out, the query you rank for may not be the query the user typed — it may be one of several sub-questions the model generated on its own. Your page has to answer a specific sub-question well, not the broad topic vaguely.
ChatGPT and Perplexity work on the same broad principle even though the plumbing differs: a retrieval step pulls candidate documents, then a generation step composes an answer and attributes claims back to the sources it used. The practical implication is the same across all three. The unit that gets cited is a claim in a passage, not a page in the abstract.
- Citation is downstream of ordinary indexing. A page that cannot be crawled, indexed, and shown with a snippet cannot be cited.
- Query fan-out means you are competing for sub-questions the model invents, not just the question the user typed.
- What gets lifted into an answer is a specific, self-contained claim — so the claim has to be findable inside your page.
The citation stack: five conditions that must all hold
I use a simple mental model with clients, because it makes the failure point obvious. Five things have to be true at once. If any one of them is false, the rest do not matter.

Figure 2: The citation stack. Most sites fail at layer four, not layer one — the page is perfectly crawlable and says nothing new.
In my experience the failure is almost never technical. Sites that have already done the basics of search engine optimization clear the first three layers routinely. They stall at layer four, because their content is a competent restatement of what forty other pages already say. There is no reason for a model to prefer them, and no claim that belongs to them.
What Google explicitly tells you to ignore
This is the part worth reading carefully, because a whole cottage industry has grown up selling the opposite. Google’s guidance includes a mythbusting section, and it is unusually direct. Below is that guidance set against the tactics I hear proposed most often.
| Commonly sold tactic | What Google’s guidance says | What to do instead |
|---|---|---|
| Publish an llms.txt file | Not used by Google Search; will neither help nor harm rankings | Fix crawlability and indexing the ordinary way |
| “Chunk” content into small AI-readable blocks | No requirement; systems understand multiple topics on a page | Structure for human readers with real headings |
| Rewrite copy in a special AI-friendly style | Not needed; systems understand synonyms and intent | Write clearly for the audience you actually have |
| Buy or seed brand “mentions” across the web | Inauthentic mentions are not as helpful as they seem | Earn coverage that a human would find credible |
| Add special schema to qualify for AI answers | No special schema.org markup is required | Keep schema for rich results; do not expect AI gains |
Google’s own summary line is worth quoting because it settles the terminology argument too: from Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO. If someone is selling you a separate GEO service that consists mainly of the tactics in the left column, you are buying a rebrand.
What actually moves the needle
Strip out the noise and the real work is narrower than most agencies would like it to be. Google’s advice reduces to one principle: create content people find unique, compelling, and useful, and it will influence your presence in generative AI search more than any other single thing.
The distinction Google draws is between commodity and non-commodity content. Commodity content is common knowledge that could have come from anyone — the example given is something like “7 Tips for First-Time Homebuyers.” Non-commodity content carries a first-hand or expert view that goes beyond what is generally known. Their example is telling: “Why We Waived the Inspection & Saved Money: A Look Inside the Sewer Line.” It is specific, it is owned by someone, and it could not be generated from general knowledge.
Here is how I translate that into an editorial checklist for clients.
- Every substantial page contains at least one claim that could only come from you — your data, your project, your measurement, your refusal.
- That claim sits in a short, self-contained passage that reads correctly when lifted out of context.
- Numbers are attributed inline, with the source and the limits of the sample stated plainly.
- The page answers a specific question rather than covering a topic, so it can match a fan-out sub-query.
- Headings describe what the section actually says, not what the section is about.
- The page is genuinely useful to a person who reads it, which is the test Google keeps returning to.
The self-contained passage point deserves emphasis, because it is the only piece of the list that is really about the AI reading experience. A model attributing a claim needs the claim to make sense on its own. A sentence that says “as we noted above, this typically runs about 30% higher” is unusable. A sentence that says “in the workflow we measured, order handling ran about 34 minutes per order” can be lifted, quoted, and credited.
Structure, and how to know whether it is working
Structural clarity still matters, but for the ordinary reason rather than a mystical one. Google’s position is that perfect semantic HTML is not required, but that clear organization, real headings, and a good page experience help both people and the systems that parse pages. That is the same advice as ever, which is why the fundamentals of organizing content into topic clusters and pillar pages have not been invalidated by any of this. If anything, fan-out rewards a cluster: several specific pages answering adjacent sub-questions is exactly the shape the retrieval step is looking for.
On measurement, be careful. Search Console now includes a Generative AI performance report showing impressions within AI features on Search and Discover, and that is the only first-party data that exists. Google is explicit that no third-party tool has access to its internal ranking or AI systems. Tools that claim to track your “AI visibility score” are inferring from samples, which can be useful directionally and should never be treated as ground truth.
The honest position to take with a client is this: you can measure impressions in Google’s AI features, you can watch referral traffic from ChatGPT and Perplexity in your analytics, and beyond that you are working on leading indicators. Publish things only you can publish, keep the technical foundation clean, and check the report monthly.
Frequently Asked Questions
Do I need an llms.txt file to be cited by AI search engines?
Not for Google. Its guidance states directly that llms.txt files and similar “special” markup are not used by Google Search, and that maintaining one will neither harm nor help your visibility or rankings there. If you want one for other services that do read it, that is fine, but it should not be sold to you as a Google AI optimization tactic.
Is GEO different from SEO?
Google’s position is that from its perspective, optimizing for generative AI search is optimizing for the search experience, and therefore still SEO. The generative features are rooted in the same core ranking and quality systems. The useful part of the GEO conversation is the emphasis on unique claims and extractable passages; the unhelpful part is the invented technical rituals.
What is query fan-out and why does it matter for my content?
Query fan-out is when the model takes a single user question and generates a set of concurrent related sub-queries to retrieve more relevant results. It matters because you are no longer competing only for the phrase someone typed. A page that answers one specific sub-question thoroughly can be retrieved for a broad question it does not directly match.
Does structured data help my pages appear in AI answers?
Google says structured data is not required for generative AI search and there is no special schema.org markup to add. It is still worth using as part of your overall SEO strategy because it makes you eligible for rich results in ordinary Search. Just do not expect schema alone to produce AI citations.
How can I tell whether AI search engines are citing me?
Use the Generative AI performance report in Google Search Console, which shows impressions within generative AI features on Search and Discover. For ChatGPT and Perplexity, watch referral sources in your analytics. Be skeptical of third-party “AI visibility” scores — Google states plainly that no third-party tool has access to its internal ranking or AI systems.
What single change makes the biggest difference?
Publish something only you can publish. A number you measured, a project you ran, a decision you made and why. Google’s guidance singles out non-commodity content with a unique point of view as more influential than any other suggestion in its guide. A page that restates common knowledge gives a model no reason to prefer it over any other source.
MEAN Consultors builds SEO and content programs around unique, citable claims — the thing Google says matters more than any technical tactic.