What is llms.txt? A complete guide with examples

·7 min read

llms.txt is a Markdown file at the root of a website, at /llms.txt, that gives AI models a short summary of the site and a curated list of links to the pages worth reading. Jeremy Howard, co-founder of Answer.AI, proposed it on September 3, 2024, and the spec lives at llmstxt.org. It is a reading list for AI agents. It does not control crawling the way robots.txt does, and Google Search ignores it.

That last point decides how much time it deserves. An llms.txt takes about an hour to write and helps the coding agents and assistants that fetch it, mostly when someone asks them about your docs or product. It will not lift your rankings or get you into AI Overviews. Write one, keep it accurate, and spend your real effort elsewhere.

The llms.txt format, with an example

An llms.txt file is plain Markdown in a fixed order. It opens with the site's name and a short summary, then lists links in sections. Here is a complete file for a made-up invoicing API.

# Ledgerline

> Ledgerline is an invoicing API for B2B software companies. It creates invoices, collects card and bank payments, and syncs paid invoices to your accounting software.

All API requests need a secret key in the Authorization header. Amounts are integers in the currency's smallest unit, so $12.50 is sent as 1250.

## Docs

- [Quickstart](https://example.com/docs/quickstart.md): Create an API key and send your first invoice
- [Invoices API](https://example.com/docs/api/invoices.md): Create, finalize, void and list invoices
- [Webhooks](https://example.com/docs/webhooks.md): Event types, retry schedule and signature checks

## Product

- [Pricing](https://example.com/pricing.md): Per-invoice fees and volume tiers
- [Integrations](https://example.com/integrations.md): Accounting and CRM sync options

## Optional

- [Changelog](https://example.com/changelog.md): API changes by date
- [Terms of service](https://example.com/legal/terms.md)

Each part has one job.

  • # Ledgerline is the H1, the site or project name, and the only required line.
  • The > blockquote is the summary. Say what the product is and who it is for, in words a model can repeat.
  • The paragraph after it holds facts an agent needs before it opens any link, like Ledgerline's amounts in cents. Any Markdown except headings works here.
  • Each ## section is a list of Markdown links, optionally followed by a colon and a note. Agents use the note to pick a link, so make it specific.
  • ## Optional holds links an agent may skip when it needs a shorter context. Changelogs and legal pages go there.

For a live file, see our own /llms.txt, which has one section per tool category.

llms.txt vs robots.txt and sitemap.xml

robots.txt decides which bots may crawl, sitemap.xml lists the URLs you want indexed, and llms.txt points a model at the few pages that explain your site.

robots.txt sitemap.xml llms.txt
Job Allow or block crawlers List every indexable URL Summarize the site and link key pages
Format Plain-text directives XML Markdown
Who reads it Search and AI crawlers Search engines Some AI agents and coding tools
Controls access Yes No No
Effect on Google Search Decides what Googlebot may fetch Helps Google find URLs None, Google ignores it

So a "do not train on this site" line in llms.txt does nothing. Allowing or blocking GPTBot and ClaudeBot happens in robots.txt alone. The robots.txt checker for AI crawlers shows which rule blocks which bot and the line to change.

What is llms-full.txt?

llms-full.txt is one file holding the full text of your docs in Markdown, so an agent can load everything in one request instead of following links. It is not in the llmstxt.org spec. Documentation platforms made it common. Mintlify generates both files for every docs site it hosts, and OpenAI links an llms-full.txt for its API docs from its own llms.txt.

It suits docs small enough to fit in a model's context window. Large docs sites do better with per-page Markdown. Our /llms-full.txt holds every tool's inputs, limits, non-goals and examples in one file.

The .md page convention

The spec asks sites to serve a clean Markdown copy of useful pages at the same URL plus .md, so /docs/webhooks gets a twin at /docs/webhooks.md. Agents find them through a rel="alternate" link with type text/markdown and get the content without navigation, scripts or cookie banners. That is why the example's links end in .md.

Some agents skip the twin and request the normal URL with an Accept: text/markdown header. A site that answers those requests with Markdown gets the same result without a second set of URLs.

Which AI tools actually read llms.txt?

Coding agents and a few AI crawlers fetch llms.txt, but no major AI search engine has said it uses the file to choose what to cite.

In June 2025, Google's John Mueller posted on Bluesky that, judging by server logs, "no AI system currently uses llms.txt." A year later Ahrefs checked May 2026 logs for 137,210 domains. About 28% had a valid llms.txt, an upper bound by Ahrefs' own account because its customers skew technical. Of those files, 97% got no requests that month. Among the fetches that did happen, GPTBot made the most of any AI bot and Claude Code came second, ahead of every AI search and assistant bot.

Anthropic, OpenAI and Perplexity each publish an llms.txt for their developer docs, so coding agents can work with their APIs. Yet OpenAI's crawler documentation says robots.txt is what controls GPTBot and OAI-SearchBot, and we found no statement from OpenAI, Anthropic, Google or Perplexity that their search crawlers use llms.txt to pick sources. The spec itself says the file is meant mainly for inference, not training. It helps an agent already working with your site, not a crawler deciding whom to cite.

Does llms.txt help SEO?

No. Google Search ignores llms.txt, including in AI Overviews and AI Mode.

Google's guide to optimizing for generative AI in Search says you need no new machine-readable files, AI text files, markup or Markdown to appear in Search. It adds that maintaining an llms.txt is fine and will neither help nor harm your visibility in Google Search, because Search ignores it.

Chrome's Lighthouse does check the file, in an llms.txt audit within its Agentic Browsing category. A missing file counts as not applicable, and only a server error fails. That is no contradiction. Lighthouse asks whether an agent can use your site, and Search asks which page should rank.

Our view is to add one if you have docs, an API or a product people ask AI tools about. It costs an hour, cannot hurt rankings, and gives agents a cleaner map than your navigation does. A local bakery can skip it. For AI citations, crawler access, server-rendered pages and specific facts matter far more, which is what how to get cited by ChatGPT covers.

Where to put llms.txt and how to keep it current

Put it at the root of your domain, at https://example.com/llms.txt, served with a 200 status as text/plain or text/markdown in UTF-8. The spec also allows a subpath, so /docs/llms.txt covers the pages under /docs. Anthropic does this for Claude Code at code.claude.com/docs/llms.txt. People also search for it as "llm txt", but agents request the lowercase /llms.txt, so name the file exactly that.

  1. Write the file or generate a draft.
  2. Upload it to your web root. In Next.js that is the public folder. On WordPress it goes in the site's root folder, unless an SEO plugin serves it for you.
  3. Fetch it with curl and confirm the first line is your # title, not HTML.
  4. Open every link and fix any that redirect, 404 or need a login.

To keep it current, generate it from the same source as your sitemap. Our /llms.txt is built from our tool registry on every deploy. It lists each public tool with its one-line promise and credit price, grouped by SEO, AEO and GEO, so a new tool shows up without anyone editing the file. If yours is hand-written, put it on the release checklist next to the sitemap.

Common llms.txt mistakes

The mistake that breaks the file outright is a /llms.txt URL that returns HTML. Single-page apps often answer unknown paths with the homepage and a 200 status, so the file seems to exist when it does not. The rest are about content.

  • Pasting the sitemap. Pick the pages you would send a new hire to.
  • Bare or relative URLs. Use [name](https://full-url) so parsers read every link.
  • Crawl rules. Disallow lines and "no training" notes belong in robots.txt.
  • Marketing copy in the summary. "The leading AI-powered platform" gives a model nothing to repeat.
  • Dead or gated links. A page behind a login, a redirect or a 404 wastes the agent's request.
  • Bot protection on the file. If your CDN challenges unknown bots, agents get a challenge page instead of Markdown.

After you publish, run the AI Agent Readiness Checker. It validates llms.txt syntax, checks that each link is an absolute public URL and tests whether your site answers Accept: text/markdown requests. A run costs 10 credits.

Generate your llms.txt from your homepage

Our llms.txt generator drafts the file from your live homepage. Enter your site URL and, if you want, the name to use as the H1. It checks whether you already publish /llms.txt and takes the site name from og:site_name, WebSite or Organization markup or your title tag, and the summary from your meta description. Then it follows up to eight same-origin links from the homepage. Docs, guides and blog pages come first, then about, pricing and product pages, with legal pages last under Optional. It skips login, account and checkout pages and takes at most two pages per site section.

Every link in the draft is a URL the run fetched with a 2xx status, titled with that page's H1 and described with its meta description. Nothing is guessed. The generator does not run JavaScript, so links your homepage renders only client-side may be missing. If you already publish a file, the result says so and tells you to compare the two before replacing anything. A run costs 15 credits.

Keep reading