LLM SEO: the two ways into a language model's answer
·7 min read
LLM SEO is the work of getting your site into the answers that ChatGPT, Claude, Gemini and Perplexity write. A language model reaches your content by two routes. Training crawlers such as GPTBot and ClaudeBot collect pages that may go into the next model, and search crawlers such as OAI-SearchBot and Claude-SearchBot index pages that the model retrieves and cites while it answers. You can only allow or block the first route, so nearly all of the practical work happens on the second.
The two routes break in different ways. Block a training crawler and today's answers do not change. Block a search crawler and your pages stop appearing as sources in that assistant's answers. The mistake to avoid is blocking OAI-SearchBot to keep content out of OpenAI's training data. That removes the site from ChatGPT search and leaves training untouched, because training is GPTBot's job.
Training vs retrieval: two routes into an LLM's answer
Training data is what a model learned before its cutoff date, and retrieval is what it looks up while it writes the answer.
| Training data | Retrieval | |
|---|---|---|
| When your page gets in | Before the model's training cutoff | Seconds before the answer |
| Who fetches it | GPTBot, ClaudeBot, CCBot | OAI-SearchBot, Claude-SearchBot, Googlebot, PerplexityBot |
| What the answer shows | Your brand or facts, usually with no link | A citation with your URL |
| What you control | Allow or block, for future models only | Access, and whether your page is found and worth quoting |
| When a change shows up | In a later model, if at all | Once the page is recrawled |
Route 1: training data and the crawlers that collect it
Training data is the text a model learned from, and your only lever is letting crawlers in or keeping them out. Anthropic's web search docs describe search as the way Claude gets information past its knowledge cutoff, so anything a model knows without searching comes from training.
GPTBot and ClaudeBot collect pages for OpenAI's and Anthropic's future models. Both vendors say a robots.txt block tells them not to train on your content. Anthropic says the block covers your future pages, so models already trained keep what they read.
CCBot is Common Crawl's crawler, and model builders train on its open web archive. In the GPT-3 paper, filtered Common Crawl made up 60% of the training mix by weight.
Google-Extended is a robots.txt token, not a crawler. Googlebot does the fetching, and the token decides whether that content may train future Gemini models and ground answers in Gemini Apps and Vertex AI. Google says it has no effect on inclusion or ranking in Google Search.
You cannot steer what a trained model says about you for a given question, and opting out only works forward. For most businesses I would leave the training crawlers allowed, because a model that read your pages can name you in answers where it never searches. Block them if your content is what you sell, such as paywalled reporting or course material. Our ClaudeBot guide walks through that call for Anthropic's three bots.
Route 2: retrieval at answer time
Retrieval is how an assistant finds pages while it writes, and it is the only route that gives you a linked citation. Each assistant runs it on a search index, and Google's own guide calls optimizing for its generative AI features still SEO.
ChatGPT
ChatGPT search rewrites the prompt into search queries and may send them to third-party providers, and OpenAI's help page for business workspaces names Bing as one. Sites that block OAI-SearchBot do not appear in ChatGPT's search answers, only as plain navigational links. So a page needs to be in Bing's index and open to OAI-SearchBot. The full setup is in how to get cited by ChatGPT.
Claude
Claude-SearchBot indexes pages for Claude's search results, and Claude-User fetches a page when a user's question needs it. Brave Search joined Anthropic's subprocessor list in March 2025, and Simon Willison showed that Claude's citations for a test query matched Brave's results. Anthropic's docs say Claude searches when a question needs current or changing information and answers from memory for stable facts.
Google AI features
AI Overviews and AI Mode draw on Google's own index. Google's AI features guide says a page needs to be indexed and eligible for a snippet, with no extra technical requirements, and both features may run several related searches per question, which Google calls query fan-out. The Search generative AI setting in Search Console, open to all sites since August 31, 2026, takes a site out of both when set to Exclude, without touching ranking or training.
Perplexity
PerplexityBot builds the index behind Perplexity's answers and is not used for model training. Perplexity-User fetches pages for a user's question and generally ignores robots.txt, since a person asked for it.
Which AI crawlers to allow for LLM SEO
Allow every search and user-fetch bot, and decide about the training bots separately.
| Bot | Company | Purpose | Route | Vendor docs |
|---|---|---|---|---|
| GPTBot | OpenAI | Model training | Training | OpenAI |
| OAI-SearchBot | OpenAI | ChatGPT search index | Retrieval | OpenAI |
| ChatGPT-User | OpenAI | Fetches pages users ask about | Retrieval | OpenAI |
| ClaudeBot | Anthropic | Model training | Training | Anthropic |
| Claude-SearchBot | Anthropic | Claude search index | Retrieval | Anthropic |
| Claude-User | Anthropic | Fetches pages users ask about | Retrieval | Anthropic |
| Google-Extended | Control token for Gemini training and grounding | Training, plus Gemini grounding | ||
| Googlebot | Google Search index | Retrieval | ||
| CCBot | Common Crawl | Open web archive | Training | Common Crawl |
| PerplexityBot | Perplexity | Perplexity search index | Retrieval | Perplexity |
| Perplexity-User | Perplexity | Fetches pages users ask about | Retrieval | Perplexity |
If you decide to keep content out of training, block only the training tokens and let everything else follow your default rules:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
User-agent: *
Disallow: /cart/Blocking Google-Extended also keeps your pages out of grounding in Gemini Apps, which is a retrieval cost. And a CDN rule that challenges unknown bots blocks search crawlers whatever the file says. Our robots.txt checker shows which rule each of these bots matches for a given path.
How language models choose what to cite
No AI company publishes how its model picks the passage it cites, so be wary of anyone selling the formula. Two things are on record. Claude's web search API returns each citation with at most 150 characters of the cited text, so a citation points at a passage, not a whole page. And the GEO paper, led by researchers at Princeton and IIT Delhi (KDD 2024), found that adding quotations, statistics and cited sources raised a source's visibility in generated answers by 30 to 40% on its main measure, using a benchmark of 10,000 queries. Keyword stuffing did about 10% worse than the unedited text on Perplexity.
The rest is our reasoning, not vendor documentation. Four traits make a retrieved passage easier to quote.
A direct answer in the first sentence. Under "How long does Acme Backup keep old versions?", open with "Acme Backup keeps 30 days of file versions on every plan." The model gets a span that matches the question.
A passage that stands alone. A 150-character citation is about one sentence. If that sentence says "as mentioned above" or starts with "This means", the quote has no subject.
Specific facts and numbers. A number with a unit and a named source is something a model can repeat with confidence. "Industry-leading uptime" is not.
Clear entity names. Write "Acme Backup" and "Google Search Console", not "our tool" and "the console", and use one product name across sections, titles and Organization schema.
This does not mean chopping pages into fragments, and Google's guide says there is no need to. Keep sections whole and let each first sentence carry its own subject. For before-and-after rewrites, see how to optimize content for AI search.
Does LLM SEO need llms.txt or special markup?
Not for Google. Its AI optimization guide says AI search needs no special text files, markup or Markdown pages, that Google Search ignores llms.txt, and that structured data is not required. It also says no third-party tool has access to Google's ranking or AI systems, so read any "AI visibility score", ours included, as a writing aid.
Check which passages on a page are easy to quote
Our Citability Analyzer reads one URL and splits its main content into sections at each h1, h2 and h3. For each section it checks whether the first sentence answers directly, meaning 50 words or fewer, not a question and not an opener like "In this article". It flags passages that lean on text outside them, such as "as mentioned above" or a first sentence starting "It is", and counts statistics, attributed quotes and links to other sites.
You get each section's lead sentence and counts, a 0 to 100 score that weights answer-first and standalone sections most, and findings that name the sections to review. It reads the HTML your server returns, so text that only appears after JavaScript runs is not checked. Pages that declare a language other than English, and pages under 150 words of prose, get no score. The tool does not rewrite anything or predict citations. A run costs 5 credits.