AI readiness for websites: what AI agents need from your pages
·8 min read
AI readiness for a website means an AI agent can read your pages and finish a task on them for a person, such as booking a table, buying a product or sending a quote request. That is the AI readiness this post covers, not whether your company is ready to adopt AI. Most of it is familiar work, like content in the server HTML, forms and buttons with real labels, stable URLs and no bot wall for agents a person sent. Newer standards such as llms.txt, Markdown responses, WebMCP and agent cards are early and optional, so they come second.
What AI readiness means for a website
Website AI readiness is about agents that act, which is a different job from the crawlers that index and cite you. Crawlers such as GPTBot and OAI-SearchBot fetch pages so a model can train on them or cite them in search. An agent shows up because one person asked it to do something. Google's Chrome auto browse, offered to AI Pro and Ultra subscribers in the U.S., can search for products, add them to a cart and apply discount codes.
Agents read a page in one of two ways. Some request the URL and parse the HTML, as a crawler does. Others drive a real browser and read the accessibility tree, the structure the browser builds for screen readers. Chrome's Lighthouse documentation says agents rely on that tree "as their primary data model." A page that works with a screen reader is most of the way to working for an agent.
Put the content in the server HTML
Agents that request a URL without opening a browser see only the HTML your server returns. If the price, stock level, opening hours or booking form appear only after JavaScript runs, those agents get an empty template. Browser agents run your scripts, but server rendering also fixes the page for the AI crawlers that decide whether you get cited.
AI Crawler View compares a page's raw HTML with the rendered version and lists the headings, prices, links and structured data that need JavaScript. The fix is server-side rendering or static generation for those pages, not a separate version for bots.
AI agent readiness starts with forms and buttons
An agent can only use a control it can identify, so every button, field and dialog needs an accessible name. Chrome announced an experimental Agentic Browsing category for Lighthouse in June 2026, and it needs Chrome 150 or later. It reports a pass ratio, not a 0 to 100 score, and its accessibility audits focus on names, labels, roles and visibility.
Here is the difference on a booking widget:
<!-- The agent finds an unnamed clickable div and a field with no label -->
<div class="btn" onclick="search()"><svg>...</svg></div>
<input placeholder="Date">
<!-- The agent finds a "Check-in date" field and a "Search" button -->
<label for="checkin">Check-in date</label>
<input id="checkin" name="checkin" type="date">
<button type="submit" aria-label="Search">
<svg aria-hidden="true">...</svg>
</button>The rules that matter most:
- Use native
<button>,<a href>,<select>and<input>elements. A<div>with a click handler has no role, so nothing tells the agent it can be clicked. - Give every field a visible
<label>tied to it withfor, plus anameattribute. A placeholder is a weak fallback that vanishes once the field has a value. - Name icon-only buttons, for example
aria-label="Close menu". - Keep IDs unique. A label that points at a repeated ID can attach to the wrong field.
- Reserve space for images, ads and cookie banners. Lighthouse counts layout shift in the same category because a button that moves after the agent reads the page can cause a misclick. The Core Web Vitals checker reports CLS from real users when field data exists.
Give every page and state a stable URL
Anything an agent might need to reach should have its own URL that loads the same way every time. An agent that can open a search URL with the query and sort order in it gets there in one step. If search, filters and booking steps live only in JavaScript state, it has to repeat every click, and each click is another chance to fail.
/search?q=oak+desk&sort=price_asc
/desks/oak-desk?size=160x80&finish=natural
/book?date=2026-11-14&time=19:30&party=4Return a real 404 for products that no longer exist. A "not found" page served with status 200 looks like success to an agent that checks the status code.
Let user-triggered agents past your bot wall
An agent working for a customer should get the pages the customer would get, so keep bot challenges off the pages people delegate. OpenAI's crawler docs say ChatGPT-User fetches pages because a user asked, and that robots.txt rules may not apply to it. The real gate for these agents is your CDN or firewall, and a rule written for scrapers can catch them too. The GEO checklist shows how to find those rules in Cloudflare and in your access logs.
Put challenges where abuse happens, such as login, sign-up and payment, and leave product pages, search and booking open. Cloudflare recognizes signed agents that prove who they are with Web Bot Auth signatures, so a rule can let verified agents through and still challenge unknown bots. Google also says auto browse pauses and asks the user before a purchase or a social post, so the person still approves the steps that matter.
Which agent-ready website standards do agents use?
Coding agents use one of them, Lighthouse audits two, and we found no major consumer assistant that says it relies on any of them. Two are still drafts.
| Standard | Proposed by | What it gives an agent | Confirmed use |
|---|---|---|---|
| llms.txt | Jeremy Howard of Answer.AI, 2024 | A Markdown map of your key pages | Lighthouse audits it. Google Search ignores it. |
| Markdown responses | Standard HTTP content negotiation. Cloudflare added a converter in February 2026 | The page without markup, in far fewer tokens | Claude Code and OpenCode send the header, per Cloudflare |
| WebMCP | Engineers at Microsoft and Google, in a W3C community group | Your forms and functions as tools with typed inputs | Chrome origin trial. No shipping agent named. |
| MCP server card | The MCP project, which Anthropic started in 2024 | Where your MCP server is and what it offers | Draft proposal |
| A2A agent card | Google, now a Linux Foundation project | What your own agent can do, for other agents | Agent-to-agent systems. No browser agent named. |
llms.txt
llms.txt is a Markdown file at /llms.txt that points agents to your most useful pages. Google's AI optimization guide says Search ignores it and that you need no AI text files, markup or Markdown versions to appear in Search. Add one if developers' agents read your docs or API. Our llms.txt guide covers the format and which tools fetch it.
Markdown responses
An agent that sends Accept: text/markdown is asking for the page without the HTML. Cloudflare's Markdown for Agents converts pages at the edge on Pro, Business and Enterprise plans, and Cloudflare says Claude Code and OpenCode already send that header. Its own launch post went from 16,180 tokens as HTML to 3,150 as Markdown. Google's guide says Search needs no Markdown, so this helps agents working within a token budget and does nothing for rankings.
GET /pricing HTTP/1.1
Host: example.com
Accept: text/markdown, text/html;q=0.9
HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
Vary: AcceptSend Vary: Accept on both versions, or a cache can hand the Markdown to a browser.
WebMCP
WebMCP lets a page declare its forms and functions as tools an agent can call with typed inputs, instead of guessing from the layout. Engineers at Microsoft and Google wrote the explainer, and it is now a draft in the W3C Web Machine Learning Community Group. Chrome opened an origin trial in Chrome 149. Its WebMCP docs list only test extensions as clients and say they are separate from Gemini in Chrome, so no shipping agent is confirmed to call these tools yet.
The declarative version is two attributes on a form you already have:
<form action="/book" toolname="book_table"
tooldescription="Book a table for a date, time and party size">
<label for="date">Date</label>
<input id="date" name="date" type="date">
<label for="party">Party size</label>
<input id="party" name="party" type="number" min="1" max="12">
<button type="submit">Book</button>
</form>Add them to a key form if it is a small change. The names may change before the API ships, and the labels help every agent today.
MCP server cards and A2A agent cards
Both are JSON files that describe a service you already run, so skip them unless you run one. Anthropic open-sourced the Model Context Protocol in November 2024 and donated it to the Linux Foundation's Agentic AI Foundation in December 2025. A server card that tells agents where your MCP server lives is still a draft, SEP-2127, and its path is not final.
A2A is Google's agent-to-agent protocol, a Linux Foundation project since June 2025. Its spec puts the agent card at /.well-known/agent-card.json. It is for agents talking to other agents and does nothing for a browser agent visiting your shop.
If you publish neither, check that those paths return 404. A catch-all route that answers them with your home page gives probing agents HTML they cannot parse.
What to fix first
Start with the pages people would hand to an agent, like pricing, product, booking and contact.
- Make their content and forms present in the server HTML.
- Give every button, field and dialog an accessible name, and stop layout shifts near controls.
- Give search results, filters and booking steps their own URLs.
- Confirm your CDN and firewall let user-triggered agents reach those pages.
- If developers use your docs or API, add llms.txt and Markdown responses.
- Add WebMCP attributes to key forms when it is cheap, and publish agent cards only for MCP servers or agents you actually run.
The first four help your human visitors as much as any agent, which is why we would do them even if agents stayed niche.
Check your site's AI agent readiness
Our AI Agent Readiness Checker runs these checks on one URL. It validates your llms.txt format and link URLs, or a draft you paste before you publish it. It checks the HTML against the accessibility rules in Lighthouse's Agentic Browsing category, such as button names, dialog names, required ARIA attributes and duplicate IDs.
It also lists every form without WebMCP toolname and tooldescription, plus fields without a name or description. It requests the page with Accept: text/markdown and checks for Vary: Accept, and it probes the MCP and A2A discovery paths. Each finding names the page attribute, llms.txt line or discovery file to fix. If bot protection answers instead of the file, the report says the check could not run instead of calling the file missing. A run costs 10 credits. It does not measure layout shift, which the Core Web Vitals checker covers.