LLM SEO: the two ways into a language model's answer

·7 min read

LLM SEO is the work of getting your site into the answers that ChatGPT, Claude, Gemini and Perplexity write. A language model reaches your content by two routes. Training crawlers such as GPTBot and ClaudeBot collect pages that may go into the next model, and search crawlers such as OAI-SearchBot and Claude-SearchBot index pages that the model retrieves and cites while it answers. You can only allow or block the first route, so nearly all of the practical work happens on the second.

The two routes break in different ways. Block a training crawler and today's answers do not change. Block a search crawler and your pages stop appearing as sources in that assistant's answers. The mistake to avoid is blocking OAI-SearchBot to keep content out of OpenAI's training data. That removes the site from ChatGPT search and leaves training untouched, because training is GPTBot's job.

Training vs retrieval: two routes into an LLM's answer

Training data is what a model learned before its cutoff date, and retrieval is what it looks up while it writes the answer.

Training data Retrieval
When your page gets in Before the model's training cutoff Seconds before the answer
Who fetches it GPTBot, ClaudeBot, CCBot OAI-SearchBot, Claude-SearchBot, Googlebot, PerplexityBot
What the answer shows Your brand or facts, usually with no link A citation with your URL
What you control Allow or block, for future models only Access, and whether your page is found and worth quoting
When a change shows up In a later model, if at all Once the page is recrawled

Route 1: training data and the crawlers that collect it

Training data is the text a model learned from, and your only lever is letting crawlers in or keeping them out. Anthropic's web search docs describe search as the way Claude gets information past its knowledge cutoff, so anything a model knows without searching comes from training.

GPTBot and ClaudeBot collect pages for OpenAI's and Anthropic's future models. Both vendors say a robots.txt block tells them not to train on your content. Anthropic says the block covers your future pages, so models already trained keep what they read.

CCBot is Common Crawl's crawler, and model builders train on its open web archive. In the GPT-3 paper, filtered Common Crawl made up 60% of the training mix by weight.

Google-Extended is a robots.txt token, not a crawler. Googlebot does the fetching, and the token decides whether that content may train future Gemini models and ground answers in Gemini Apps and Vertex AI. Google says it has no effect on inclusion or ranking in Google Search.

You cannot steer what a trained model says about you for a given question, and opting out only works forward. For most businesses I would leave the training crawlers allowed, because a model that read your pages can name you in answers where it never searches. Block them if your content is what you sell, such as paywalled reporting or course material. Our ClaudeBot guide walks through that call for Anthropic's three bots.

Route 2: retrieval at answer time

Retrieval is how an assistant finds pages while it writes, and it is the only route that gives you a linked citation. Each assistant runs it on a search index, and Google's own guide calls optimizing for its generative AI features still SEO.

ChatGPT

ChatGPT search rewrites the prompt into search queries and may send them to third-party providers, and OpenAI's help page for business workspaces names Bing as one. Sites that block OAI-SearchBot do not appear in ChatGPT's search answers, only as plain navigational links. So a page needs to be in Bing's index and open to OAI-SearchBot. The full setup is in how to get cited by ChatGPT.

Claude

Claude-SearchBot indexes pages for Claude's search results, and Claude-User fetches a page when a user's question needs it. Brave Search joined Anthropic's subprocessor list in March 2025, and Simon Willison showed that Claude's citations for a test query matched Brave's results. Anthropic's docs say Claude searches when a question needs current or changing information and answers from memory for stable facts.

Google AI features

AI Overviews and AI Mode draw on Google's own index. Google's AI features guide says a page needs to be indexed and eligible for a snippet, with no extra technical requirements, and both features may run several related searches per question, which Google calls query fan-out. The Search generative AI setting in Search Console, open to all sites since August 31, 2026, takes a site out of both when set to Exclude, without touching ranking or training.

Perplexity

PerplexityBot builds the index behind Perplexity's answers and is not used for model training. Perplexity-User fetches pages for a user's question and generally ignores robots.txt, since a person asked for it.

Which AI crawlers to allow for LLM SEO

Allow every search and user-fetch bot, and decide about the training bots separately.

Bot Company Purpose Route Vendor docs
GPTBot OpenAI Model training Training OpenAI
OAI-SearchBot OpenAI ChatGPT search index Retrieval OpenAI
ChatGPT-User OpenAI Fetches pages users ask about Retrieval OpenAI
ClaudeBot Anthropic Model training Training Anthropic
Claude-SearchBot Anthropic Claude search index Retrieval Anthropic
Claude-User Anthropic Fetches pages users ask about Retrieval Anthropic
Google-Extended Google Control token for Gemini training and grounding Training, plus Gemini grounding Google
Googlebot Google Google Search index Retrieval Google
CCBot Common Crawl Open web archive Training Common Crawl
PerplexityBot Perplexity Perplexity search index Retrieval Perplexity
Perplexity-User Perplexity Fetches pages users ask about Retrieval Perplexity

If you decide to keep content out of training, block only the training tokens and let everything else follow your default rules:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /cart/

Blocking Google-Extended also keeps your pages out of grounding in Gemini Apps, which is a retrieval cost. And a CDN rule that challenges unknown bots blocks search crawlers whatever the file says. Our robots.txt checker shows which rule each of these bots matches for a given path.

How language models choose what to cite

No AI company publishes how its model picks the passage it cites, so be wary of anyone selling the formula. Two things are on record. Claude's web search API returns each citation with at most 150 characters of the cited text, so a citation points at a passage, not a whole page. And the GEO paper, led by researchers at Princeton and IIT Delhi (KDD 2024), found that adding quotations, statistics and cited sources raised a source's visibility in generated answers by 30 to 40% on its main measure, using a benchmark of 10,000 queries. Keyword stuffing did about 10% worse than the unedited text on Perplexity.

The rest is our reasoning, not vendor documentation. Four traits make a retrieved passage easier to quote.

A direct answer in the first sentence. Under "How long does Acme Backup keep old versions?", open with "Acme Backup keeps 30 days of file versions on every plan." The model gets a span that matches the question.

A passage that stands alone. A 150-character citation is about one sentence. If that sentence says "as mentioned above" or starts with "This means", the quote has no subject.

Specific facts and numbers. A number with a unit and a named source is something a model can repeat with confidence. "Industry-leading uptime" is not.

Clear entity names. Write "Acme Backup" and "Google Search Console", not "our tool" and "the console", and use one product name across sections, titles and Organization schema.

This does not mean chopping pages into fragments, and Google's guide says there is no need to. Keep sections whole and let each first sentence carry its own subject. For before-and-after rewrites, see how to optimize content for AI search.

Does LLM SEO need llms.txt or special markup?

Not for Google. Its AI optimization guide says AI search needs no special text files, markup or Markdown pages, that Google Search ignores llms.txt, and that structured data is not required. It also says no third-party tool has access to Google's ranking or AI systems, so read any "AI visibility score", ours included, as a writing aid.

Check which passages on a page are easy to quote

Our Citability Analyzer reads one URL and splits its main content into sections at each h1, h2 and h3. For each section it checks whether the first sentence answers directly, meaning 50 words or fewer, not a question and not an opener like "In this article". It flags passages that lean on text outside them, such as "as mentioned above" or a first sentence starting "It is", and counts statistics, attributed quotes and links to other sites.

You get each section's lead sentence and counts, a 0 to 100 score that weights answer-first and standalone sections most, and findings that name the sections to review. It reads the HTML your server returns, so text that only appears after JavaScript runs is not checked. Pages that declare a language other than English, and pages under 150 words of prose, get no score. The tool does not rewrite anything or predict citations. A run costs 5 credits.

Keep reading