Markdown vs HTML: what AI crawlers read best
·7 min read
In the markdown vs html debate, HTML is what search and AI answer crawlers read, and I could find no crawler documentation or published test showing that Google, ChatGPT, Perplexity or Claude rank or cite a Markdown copy of a page over its HTML. Google's own guidance says you need no Markdown to appear in Google Search or its AI features. Markdown does help one group of readers, the AI agents that request pages with an Accept: text/markdown header, which today are mostly coding agents such as Claude Code, Cursor and OpenCode. Keep server-rendered HTML as the version everyone gets, and add Markdown through content negotiation only if agents read your docs.
If you run a shop or a local business site, this is low on your list. Fix your HTML first.
Markdown vs HTML at a glance
| HTML | Markdown | |
|---|---|---|
| Who reads it | Browsers and every search and AI crawler | Agents that ask for it, today mostly coding agents |
| What it carries | Content plus the <head>: title, canonical, meta robots, hreflang, JSON-LD |
Headings, text, links, lists and tables. No <head> at all |
| Size | Cloudflare's launch post was 16,180 tokens as HTML | The same post was 3,150 tokens as Markdown |
| Google rankings and AI features | It is what Google indexes | Google says you don't need it to appear |
| How it usually goes wrong | Content that only exists after JavaScript runs | A copy that drifts from the HTML, or a cache that serves it to browsers |
Do AI crawlers prefer Markdown?
Not that anyone has shown. Google, the one search company that has written about it, says you don't need Markdown at all. Its guide to generative AI in Search says you don't need to create "machine readable files, AI text files, markup, or Markdown" to appear in Google Search. The AI features and your website page says the same for AI Overviews and AI Mode without naming Markdown, and adds that no special schema.org markup is required either.
OpenAI's crawler documentation describes OAI-SearchBot, GPTBot and ChatGPT-User without a word about Markdown. The docs site itself does offer a Markdown copy of every page if you add .md to the URL. That is a developer docs site serving developers' agents, which is exactly where Markdown pays off.
Google's John Mueller has been blunter. A developer on Reddit proposed detecting GPTBot and ClaudeBot by user agent and serving them Markdown instead of the React page. Mueller asked whether bots would even treat a Markdown file as more than plain text, then dismissed the idea on Bluesky, as Search Engine Journal reported in February 2026.
There is also a practical reason. A Markdown file has no place for a canonical tag, a meta robots rule or JSON-LD. A crawler that saw only your Markdown would know less about the page.
Which AI agents send Accept: text/markdown?
Coding agents, mainly. In Checkly's February 2026 test, three agents put text/markdown at the front of their Accept header and four did not.
| Agent and version tested | Accept header | Prefers Markdown |
|---|---|---|
| Claude Code 2.1.38 | text/markdown, text/html, */* |
Yes |
| Cursor 2.4.28 | text/markdown,text/html;q=0.9,... |
Yes |
| OpenCode 1.2.5 | text/markdown;q=1.0, text/x-markdown;q=0.9,... |
Yes |
| OpenAI Codex | Starts with text/html |
No |
| GitHub Copilot | Starts with text/html |
No |
| Gemini CLI 0.28.2 | */* |
No |
| Windsurf | */* |
No |
I pointed a current Claude Code build (2.1.295) at a header echo page in October 2026, and it still sends text/markdown, text/html, */*.
Keep two limits in mind. These are fetches a developer triggers while working, not crawls that build a search index, so Markdown changes what the agent reads and nothing about where you rank. And the test did not cover ChatGPT, the Gemini app or Perplexity. I found no documentation saying they ask for Markdown, so don't assume it.
Markdown for LLMs: what it saves and who should bother
Markdown saves tokens, usually most of them. Cloudflare measured its own launch post at 16,180 tokens as HTML and 3,150 as Markdown, about 80% fewer. The difference is markup a model has no use for, like class names, inline scripts and SVG, navigation, footers and cookie banners. An agent with a fixed context budget can read more of your content, or more of your pages, before it runs out.
Serving your own Markdown also lets you choose what the agent reads, instead of its converter guessing which <div> holds the article.
Serve Markdown if developers point agents at your docs, API reference or help center. That is where the requests come from today. A shop, a local business or a content site gains little from it right now and should spend the time on its HTML.
How to serve Markdown to AI agents
Serve it from the same URL through content negotiation. A request whose Accept header prefers text/markdown gets Markdown, and every other request gets the HTML it always got.
GET /docs/webhooks HTTP/1.1
Host: example.com
Accept: text/markdown, text/html, */*
HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
Vary: AcceptWith Cloudflare
If your site sits behind Cloudflare, Markdown for Agents does the conversion at the edge. Cloudflare launched it on February 12, 2026. It is a toggle on Pro, Business and Enterprise plans at no extra cost. When a request's Accept header includes text/markdown, Cloudflare fetches your HTML, converts it, and returns Markdown with Vary: Accept and an x-markdown-tokens estimate. Its docs list the limits. It converts only HTML, up to 6 MiB, and it drops ETag and Last-Modified from converted responses.
Check one default before you switch it on. Converted responses carry Content-Signal: ai-train=yes, search=yes, ai-input=yes unless your origin sends its own Content-Signal header. Set one if you don't want to signal consent to AI training.
On your own server
- Build the Markdown from the same source as the HTML, such as your CMS or MDX files. Never maintain a hand-edited copy for agents.
- Return Markdown when the Accept header prefers
text/markdown, and HTML otherwise. - Send
Vary: Accepton both responses. - Test both versions with curl.
In Express, req.accepts does the negotiation. It returns HTML for browsers, for a bare */* and for requests with no Accept header.
app.get("/docs/:slug", async (req, res) => {
const page = await getDoc(req.params.slug); // one source for both formats
res.vary("Accept");
if (req.accepts(["text/html", "text/markdown"]) === "text/markdown") {
res.type("text/markdown; charset=utf-8").send(page.markdown);
} else {
res.type("html").send(page.html);
}
});# Ask for Markdown and print only the response headers
curl -s -o /dev/null -D - -H "Accept: text/markdown" https://example.com/docs/webhooks
# Ask like a browser, which should still get HTML
curl -s -o /dev/null -D - -H "Accept: text/html" https://example.com/docs/webhooksThe llms.txt proposal also suggests a Markdown twin of each page at the same URL plus .md. If you publish twins, link each one from the HTML page's head.
<link rel="alternate" type="text/markdown" href="https://example.com/docs/webhooks.md">llms.txt itself is a different file, a Markdown index of your key pages at /llms.txt that Google Search ignores. Our guides explain what llms.txt is and how to create and publish one.
The risks: duplicate URLs, cloaking and caching
Each one has a plain fix, and content negotiation on a single URL avoids the first entirely.
Duplicate URLs
A .md twin is a second URL with the same content. Point it back at the HTML page with a canonical in the HTTP response, which Google supports for non-HTML documents.
HTTP/1.1 200 OK
Content-Type: text/markdown; charset=utf-8
Link: <https://example.com/docs/webhooks>; rel="canonical"Cloaking
Google's spam policies define cloaking as showing search engines different content from what people see, to manipulate rankings and mislead users. A Markdown rendering of the same words is a format change. The trouble starts when the versions differ, such as an agent copy with extra keywords, claims or prices the page doesn't show. Switching on user agent is riskier than switching on the Accept header, because you are choosing what bots see instead of answering what they asked for.
Caching
Without Vary: Accept, a CDN or browser cache can store the Markdown and serve it to the next visitor, or hand the HTML to the agent. Browsers send long Accept strings that differ by browser and version, so a cache keyed on the raw header splits into many copies. Normalize it to two values, Markdown or HTML, before it reaches the cache key. Also check your bot protection. A WAF that challenges unfamiliar clients answers agents with a challenge page, whatever format they asked for.
Clean HTML is still the baseline
Every Markdown version starts as HTML, so good Markdown needs good HTML. Cloudflare converts the HTML your server returns, which means a page that builds its content in the browser converts to a near-empty file. Crawlers that don't run JavaScript get the same empty shell. Our guide to JavaScript SEO covers the fixes, and you can compare the raw and rendered source of any page to see what a crawler gets.
What converts well, and what crawlers parse well, is the same list:
- Real text in the initial HTML response.
- Headings in order, because
<h2>and<h3>become##and###. - Tables built with
<table>,<th>and<td>. A grid of divs converts to loose lines with no columns. - Lists as
<ul>and<ol>, links with descriptive text, and images with alt text. - The
<head>metadata Markdown cannot carry, like the title, canonical, meta robots and structured data.
Our AI readiness guide covers what else agents need from a page, such as labeled buttons and forms they can use.
Check whether your site answers Accept: text/markdown
The AI Agent Readiness Checker requests your page with Accept: text/markdown, text/html;q=0.9, */*;q=0.8 and reports the status and content type that came back, whether the response sends Vary: Accept, the x-markdown-tokens count when your server or CDN sends one, and any <link rel="alternate" type="text/markdown"> in the page head. It flags a page that errors when asked for Markdown and Markdown served without Vary: Accept. If bot protection or a rate limit answers instead, it says the check could not run rather than calling Markdown missing. It does not read the Markdown body or compare it with your HTML, so check that by hand.
The same run validates your /llms.txt, or a draft you paste, runs the checks from Lighthouse's Agentic Browsing category without a browser, covering accessibility-tree rules and WebMCP form coverage, and looks for MCP and A2A agent cards. Each finding names the header, page attribute or file to fix. A run costs 10 credits.