There is no ranking in AI search. There are four different markets that disagree with each other.
own study, 15 august 2026, method disclosed
LLM SEO — also sold as answer engine optimization, AEO or generative engine optimization — is the work of getting your pages used as a source when an assistant answers a question. The win is not a position and a click. The win is a citation inside somebody else's answer, on a page you do not control.
Almost everything written about this is advice without measurement. So I measured. Four engines, 40 real buying questions, every cited source recorded: 1,121 citations, 699 domains. The engines overlap by 4.3 to 14.0 percent — see section 02.
The finding with the most practical weight is in section 03. Of the 82 answers where the engine showed its work, not one searched the question it was asked. They rewrite it — 2.67 queries on average, up to eight — and ChatGPT then runs verification queries against the candidate's own website. That is where most sites fall out.
How this was measured
There is no official statistic for what AI assistants cite. There are vendor blogs, and there are tools that sell you a score. So I ran the thing itself.
40 questions, in five groups of eight, all of them questions a real buyer types before spending money: who to hire, what it costs, which tool to use, how to do something, and one thing versus another. They sit in the service-business space — agencies, video, branding, websites, social, SEO — because that is the market I can check the answers against.
Four engines, web search enabled, US market, one live call per engine per question on 15 August 2026: ChatGPT (gpt-5.4-mini), Perplexity (sonar-pro), Claude (claude-sonnet-4-5), Gemini (gemini-3.5-flash). That is 160 answers, with every annotated source recorded. Then Google's own organic top 20 for the same 40 questions as a control — 37 of 40 returned a result set.
Then a second pass on the infrastructure: all 902 domains that any engine cited or Google ranked were checked for robots.txt and llms.txt — and, as a control group, 980 ordinary German local-business websites (hairdressers, bakeries, workshops) collected for an earlier study, checked the same way on the same day.
What these numbers do not say
Four limits, which I would rather state myself. First: this is one snapshot on one day, in one market. AI answers are not stable — ask again next week and the sources move. Second: one model per engine. A bigger or smaller model in the same family may cite differently, and the assistant your buyers actually use may not be the one tested here.
Third: a citation is not a recommendation. Being named as a source is not the same as being named as the answer, and this study counts sources, not praise. Fourth: an engine can lean on something it never annotates. What is measured here is what each engine was willing to show as its source — that is the only part anybody outside the lab can verify.
What the numbers do say: how wide the disagreement between engines is, what kind of page gets used, and how much of it Google's ranking already predicts. That is enough to decide what to do on Monday.
Four engines, four different worlds
Start with volume, because the engines are not remotely comparable. Asked the same question, one gives you a shortlist of two sources and another gives you eighteen.
That single chart reframes the whole exercise. If your buyers use ChatGPT, you are competing for roughly two slots. If they use Perplexity, there are eighteen, and being one of them means much less. Any vendor selling you one "AI visibility score" is averaging over four markets with different rules.
Now the disagreement. For each question I compared which domains each pair of engines cited:
Read the bottom row again: the two assistants most people actually use agree on 4.3 percent of their sources. Across 31 questions where all four engines cited something, the mean number of domains shared by all four was 0.26 — and only four questions produced even one such domain.
The long tail says the same thing from the other side. Of 699 domains, 532 — 76.1 percent — were cited exactly once in the entire study. The ten most-cited domains together account for just 14 percent of all citations. There is no small club that owns AI answers in this category.
This is the good news, and it is worth stating plainly, because most coverage of AI search is written as if the door were closing. A market where three quarters of cited sources appear once is a market with room in it. It is also a market where nobody can promise you a position.
Nobody searches your question
Two of the four engines expose the search queries they ran before answering. Across 82 answers where those queries were visible, here is the count that matters:
Instead they take the question apart and rebuild it — 2.67 queries per answer on average, up to eight. Ask "Which agency should I hire to produce a corporate explainer video?" and ChatGPT goes looking for best corporate explainer video agency 2026 explainer video production agency corporate clients, then corporate explainer video agency portfolio motion graphics corporate video agency.
That alone would only be a keyword lesson. The second half is the one nobody seems to have written down.
The engine finds candidates from third-party sources, then walks to each candidate's own site to check them. In other runs it does this with an explicit site: operator — site:sculpt.co B2B social media marketing LinkedIn official. The word official shows up again and again: it is looking for the company's own statement of what it does.
This is the step where most websites lose. You can be mentioned in a listicle, land on the candidate list, get visited — and then fail the check, because the page the engine lands on says "we deliver bespoke solutions that drive impact" instead of what you actually do, for whom, at what price, with what result.
What that means for the page you own
The verification query is a fact-matching operation. It succeeds when your page states, in plain text, the things the query asks about: the service in the words buyers use, the segment you serve, the platform or method named, evidence, and — the most avoided of all — the price. Three of the five question groups in this study (price, tool choice, comparison) are answered almost entirely from pages that state specifics.
There is a second-order effect worth planning for: because the engine writes its own queries, it also writes queries you never targeted. You cannot rank for a rewrite you have not seen. What you can do is make sure the page carries the underlying facts in several plain formulations rather than one polished slogan.
Could a machine cite this sentence?
Paste a paragraph from your website. This counts the things a verification query can actually match — numbers, prices, dates, named places and methods — and the empty phrases that match nothing. It runs entirely in your browser; nothing is sent anywhere.
What kind of page gets cited
The received wisdom is that AI answers are stitched together from Reddit threads and review directories, and that a normal business site has no chance. Sorted by category, the 1,121 citations look like this:
Reddit is the single most-cited domain — and it is still under four percent of everything. The picture is not a handful of gatekeepers; it is a very long tail of ordinary websites, most of them named once.
The one place where the directories do concentrate is exactly where you would expect: "who should I hire" questions. There, directories take 14.0 percent of citations, four times their share elsewhere. On price questions it drops to 3.8 percent, on comparisons to 1.6 percent.
So the practical split is: for vendor-choice questions, a profile on the two or three directories in your category is worth having, because that is where candidate lists get built. For everything else — cost, method, comparison, how-to — the citation goes to whoever actually wrote the specifics down, and that can be you.
How much of this Google already decides
For the same 40 questions I pulled Google's organic results, then compared them against what the engines cited.
Two numbers, and they point in opposite directions. Averaged per question, 27.7 percent of the domains cited by AI also appear in Google's top 10. So seven of ten AI citations come from outside the first page. But looked at the other way, 67.2 percent of Google's top 10 gets cited by at least one engine — and there was not a single question with zero overlap.
Both are true and both matter. Ranking on page one makes you a likely candidate for citation. It is nowhere near sufficient, and it is not the only door: most of what gets cited was not on page one at all.
And the number that should end the "AI versus Google" framing entirely: Google returned an AI Overview on 37 of the 37 questions that produced results. Every single one. The biggest generative answer engine in this study was not ChatGPT — it was the search engine that already sends most of your traffic.
Which is why I am careful with the traffic argument. AI assistants still send well under one percent of referral traffic to publishers, while Google sends around 88 percent of search referrals. If someone tells you to rebuild your site because AI traffic is about to replace search traffic, the numbers do not support them. The reason to care is not the traffic AI sends. It is whether you exist in the answer at the moment somebody decides who to call.
Can the engine even read your site?
Every domain in this study got a second check: what its robots.txt says to AI crawlers, and whether it publishes an llms.txt. Then the same check on a control group — 980 ordinary German local-business websites, the kind of company that is not thinking about AI search at all.
The headline difference is not that one group blocks and the other does not. It is that one group has decided anything at all. Cited sites name GPTBot in their robots.txt eleven times as often, and publish an llms.txt more than four times as often.
Now the part that costs ordinary businesses money. OpenAI runs more than one crawler, and they do different jobs. From OpenAI's own documentation: GPTBot crawls content that may be used to train models. OAI-SearchBot is the one that, in OpenAI's words, is "used to surface websites in search results in ChatGPT's search features". ChatGPT-User is a fetch triggered by a person.
Read those two pairs together. Sites that get cited block the training crawler more often and the search crawler less often. Ordinary business sites do the exact opposite: they are twice as likely to block the one crawler that decides whether they appear in ChatGPT's search results at all.
Nobody chose that. It is what a blanket Disallow from a security plugin, a hosting default or a copied template does when it meets a crawler nobody in the company has heard of. Two other numbers from the same scan point the same way: 18.9 percent of the local-business sites have no robots.txt at all, and 14.4 percent did not answer within twelve seconds.
About llms.txt, honestly
llms.txt is a markdown file at your site root that hands an AI agent a clean summary and a map of your important pages. Jeremy Howard proposed it in September 2024; version 2 was published on 10 August 2026, OpenAI, Anthropic and Google all publish one for their own documentation, and Chrome's Lighthouse now audits for it.
And no engine has confirmed that it affects citation. So I will not tell you it is a ranking factor — that claim is currently unsupported. What the data shows is narrower and still interesting: the sites that get cited are four and a half times more likely to have one. That is a correlation with the kind of team that also does the other things right, not proof that the file did the work. It costs an hour. It cannot hurt. That is the whole case for it.
The mistake I found on my own site
This site is built for exactly what this article describes. AI crawlers are explicitly invited in robots.txt. There is an llms.txt. There is FAQPage markup on the articles. I have been writing about generative search since last year.
On 15 August 2026 I ran the check on heidarrudyi.com/cases — the page that holds every piece of work I have to show. It contained no static link to a single case study. The whole grid was drawn from a JavaScript array at load time.
Google renders JavaScript, so Google saw the portfolio. The AI crawlers I had invited in largely do not render JavaScript. They saw an empty page. A site optimized for AI visibility was hiding its fifteen pieces of proof from the machines it had personally invited.
The fix took twenty minutes: a visible text index of all fifteen links underneath the grid — useful to people as quick navigation, and real links for anything that cannot run JavaScript. Not hidden text, because hidden text is a different problem.
curl -s https://yoursite.com/page | grep -o 'href="[^"]*"' | head
I mention this because it is the most common version of the problem, and it has nothing to do with llms.txt, schema markup or any AI-specific tactic. The content simply was not there. Before optimizing anything for AI, check that a machine can see what you already have.
What I would actually do on Monday
Step four: pick your two directories, then stop. Directories take 14 percent of citations on „who should I hire" questions and almost nothing everywhere else. A profile on the two that matter in your category is worth an afternoon. A profile on twenty is a subscription you will not notice paying.
Step five: publish the specifics nobody else will. Three of the five question groups here — price, tool choice, comparison — are answered from pages that state numbers. Most companies in most industries refuse to publish prices, methods and outcomes, which means whoever does becomes the only citable source in the category. That is also why my own prices sit openly on one page rather than behind a call.
What I would not do: buy an „AI visibility score". This study measured four engines that overlap by four to fourteen percent, on one day, with results that move. A single number averaged across them tells you nothing you can act on.
Want to know what the engines say about you?
Send me your domain and the question your buyers actually ask before they hire someone like you. You get one page back with three things: what the four engines answer to that question today and who they cite, what a crawler sees when it fetches your site without JavaScript, and the specific sentences missing from your pages that the verification queries were looking for. No call required, and if nothing comes of it you keep the analysis.
Get the analysis →LLM SEO — the questions worth asking before anyone sells you a score.
What is LLM SEO?
LLM SEO — also called answer engine optimization or generative engine optimization — is the work of getting your pages used as a source when an AI assistant answers a question. It differs from classic SEO in what counts as a win: not a ranking position and a click, but a citation inside somebody else's answer. In this study of 160 AI answers, 83.1 percent of citations went to ordinary company websites rather than directories, forums or big platforms — so the target is reachable for a normal business site.
How do I get my website cited by ChatGPT?
Three things follow from the data. First, let the right crawler in: blocking GPTBot stops training, but blocking OAI-SearchBot is what removes you from ChatGPT's search results — and ordinary local business sites block the search crawler twice as often as sites that actually get cited (3.2 percent versus 1.5 percent). Second, make the claim checkable on your own domain, because ChatGPT runs verification queries like site:youragency.com … official against the candidate's own website. Third, publish the specific numbers, prices and processes those queries are looking for — an engine cannot cite what your page does not state.
Do the different AI engines cite the same sources?
No, and the gap is wider than expected. Across the same 40 questions, the pairwise overlap in cited domains was 4.3 percent (ChatGPT and Claude) to 14.0 percent (Gemini and Perplexity). Of 31 comparable questions, only four had a single domain that all four engines named. 76.1 percent of the 699 domains were cited exactly once. There is no single AI ranking to optimize for — there are four different markets.
How many sources does an AI answer actually use?
It depends entirely on the engine. Median sources per answer: Perplexity 18, Gemini 7, Claude 4, ChatGPT 2. And they often cite nothing at all — Claude gave 15 of 40 answers with no source, Gemini 13, ChatGPT 10, while Perplexity always cited something. If you are being measured on citations, which assistant your buyers use matters more than any single thing you do to your site.
Does classic SEO still help with AI search?
It helps, but it is not the same job. Averaged per question, 27.7 percent of the domains cited by AI also sat in Google's top 10, and 67.2 percent of Google's top 10 was cited by at least one engine. So ranking well makes you a candidate — but roughly seven in ten AI citations came from outside the top 10. Also worth noting: Google showed an AI Overview on 37 of the 37 questions that returned results. The largest AI answer engine is still Google.
What is llms.txt and does it work?
llms.txt is a proposed markdown file at your site root that gives AI agents a clean summary and links; Jeremy Howard published it in September 2024 and a v2 in August 2026, and Chrome's Lighthouse now audits for it. No engine has confirmed that it affects citation, so nobody should sell it as a ranking factor. What the data does show is a striking split in who bothers: 43.8 percent of the domains that got cited or ranked have one, against 9.4 percent of ordinary local business sites.
Which AI crawlers should I allow in robots.txt?
Decide per crawler, because they do different jobs. From OpenAI's own documentation: GPTBot crawls content that may be used to train models, OAI-SearchBot is the one that surfaces sites in ChatGPT search, and ChatGPT-User is a user-initiated fetch. Blocking GPTBot while allowing OAI-SearchBot is a coherent position — stay out of training, stay in the answers. Blocking everything with a blanket rule is usually not a decision at all, it is a default someone inherited from a plugin or a host.
Why can't AI see my site even though Google can?
Usually because the content is drawn by JavaScript. Google renders JS; most AI crawlers largely do not. We found this on our own site on 15 August 2026: the /cases page had no static link to any case study at all — the grid was built from a JS array, so crawlers we had explicitly invited saw an empty portfolio. The one-line check is curl -s yoursite.com/page | grep -o 'href="[^"]*"' — if the links are not in that output, they do not exist for the machine.
How was this study run?
On 15 August 2026 we put the same 40 buying-stage questions to ChatGPT (gpt-5.4-mini), Perplexity (sonar-pro), Claude (claude-sonnet-4-5) and Gemini (gemini-3.5-flash), all with web search enabled, US market, through the DataForSEO AI-optimization API — 160 answers, every cited source recorded. Google's own top 20 was pulled for the same questions as a control. Then all 902 cited or ranked domains, plus a control group of 980 ordinary German local-business sites, were checked for robots.txt and llms.txt. It is one snapshot, one market, one model per engine — a citation is not a recommendation, and an engine can use a source without annotating it.