View rendered source vs raw HTML: what crawlers see on your page

·7 min read

Raw HTML is the file your server sends. Rendered HTML is what that file becomes after the browser runs your JavaScript, and to view rendered source you need DevTools, a copy of the DOM or Search Console, because View Page Source only ever shows the raw file. The gap decides what crawlers read. Googlebot renders pages before indexing them, while GPTBot, ClaudeBot and PerplexityBot read the raw HTML only, so anything your scripts add never reaches them.

On a server-rendered site the two versions usually differ in ways nobody cares about. On a client-rendered app the raw file can be a few hundred bytes with one empty div, and that empty div is the page every AI crawler gets.

Raw HTML vs rendered HTML: what's the difference

Raw HTML is the response body exactly as the server sent it, and rendered HTML is the DOM, the browser's live model of the page, written back out as HTML after parsing and scripts.

The parser changes things first. It adds a missing <tbody>, closes tags you left open and moves anything that can't live in <head> into <body>. Then scripts run, and a framework such as React or Vue can build the entire page from an empty container:

<!-- raw HTML from the server -->
<body>
  <div id="root"></div>
  <script src="/assets/index-8f3a21c4.js"></script>
</body>

Once that script runs, the same div holds the H1, the copy, the prices and the links. A crawler that executes JavaScript sees all of it. A crawler that doesn't sees an empty div and a script tag. Tag managers do the same thing on a smaller scale when they inject JSON-LD, rewrite a title or add a canonical after load.

Which crawlers render JavaScript?

Googlebot does, and most AI crawlers don't. Google says Googlebot queues every page that returns a 200 for rendering, runs it in an evergreen version of Chromium and uses the rendered HTML to index it. The exception is a page whose original HTML or headers say noindex, where Google may skip rendering altogether. Our post on what Googlebot is and how it crawls covers the rest of that pipeline.

Crawler Runs your JavaScript? Source
Googlebot Yes, in an evergreen Chromium Google Search Central
Gemini Yes, through Googlebot's infrastructure Vercel and MERJ
Applebot May render pages in a browser Apple
GPTBot, OAI-SearchBot, ChatGPT-User No Vercel and MERJ
ClaudeBot No Vercel and MERJ
PerplexityBot No Vercel and MERJ
Meta-ExternalAgent, Bytespider No Vercel and MERJ
CCBot (Common Crawl) No Vercel and MERJ

The AI rows come from one large test of real crawler traffic, which Vercel and MERJ published in December 2024. They saw OpenAI's and Anthropic's crawlers download JavaScript files without executing them. That data is nearly two years old, and OpenAI's own crawler documentation says nothing either way about rendering. I assume raw HTML only for every AI crawler until a site's logs prove otherwise.

In the logs, a bot downloading your .js bundle is not evidence of rendering. A bot calling the API endpoint that bundle fetches is, because only a bot that runs the script would make that request. Log file analysis shows how to pull those lines out, and the AI crawler list has every user agent to filter on.

How to view page source (the raw HTML)

Press Ctrl+U in Chrome, Edge or Firefox, or Option+Command+U in Chrome on a Mac. Typing view-source: in front of the URL does the same. The tab shows the response body as plain text. Press Ctrl+F and search for a sentence from the middle of your page. If it isn't there, crawlers that skip JavaScript don't have it either.

View-source has one blind spot. It is your browser's request, with your cookies, your login and your location, so it can show a version no crawler ever gets. curl with a crawler's user agent is closer to the real thing:

curl -sL -o raw.html -w "%{http_code}\n" \
  -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
  https://example.com/pricing

That string is the one OpenAI publishes for GPTBot. A 200 means you got the page. A 403 or a challenge page can mean your CDN blocks GPTBot by name, or that it checked the IP address and caught a fake, since the real GPTBot only comes from addresses OpenAI lists. Run the same command with a browser user agent to tell the two apart.

View page source vs inspect element

View Page Source shows the file the server sent, and Inspect Element shows the live DOM. Inspect opens the DevTools Elements panel, which includes every repair the parser made and every node your scripts added, moved or deleted. That is how a tag can sit right there in Inspect Element and be missing from the source.

Two things in the Elements panel catch people out. Browser extensions inject their own markup, so a password manager or grammar checker can add nodes no crawler will ever see. The parser's repairs can also expose a real bug. If the source has a <div> or an <img> inside <head>, the browser ends the head at that element, and Elements shows every tag after it inside <body>, your meta robots and hreflang included. Google reads the head the same way. Its page metadata guidelines say it ignores any head element that comes after an invalid one.

How to view rendered source

Copy the DOM out of the browser once the page has finished loading. Any of these three ways works.

  1. In DevTools, right-click the <html> node in the Elements panel, choose Copy > Copy outerHTML and paste it into rendered.html.
  2. In the DevTools console, run copy(document.documentElement.outerHTML). The copy() helper puts the whole rendered page on your clipboard, minus the doctype, which doesn't matter here.
  3. From a terminal, run headless Chrome with --dump-dom, which prints the serialized DOM after scripts have run. Chrome's documentation points out that this output differs from what curl returns, which is exactly why you want it.
chrome --headless --dump-dom https://example.com/pricing > rendered.html

On Linux the binary is usually google-chrome. On a Mac, call the executable inside Google Chrome.app. Some bot protection turns headless Chrome away, and if you get a challenge page back, copy the DOM from DevTools instead.

For a visual check without any files, open the DevTools command menu with Ctrl+Shift+P, or Command+Shift+P on a Mac, run Disable JavaScript and reload. Whatever disappears is what GPTBot never sees.

See the rendered HTML Google indexed

Use URL Inspection in Search Console, which shows the page as Google's own renderer produced it. Google's JavaScript troubleshooting guide names it as the place to check that the rendered HTML has the content you expect.

  1. Paste the URL into the inspection bar at the top of Search Console.
  2. Click View crawled page to see the HTML of the indexed version. This view has no screenshot.
  3. For the current version, click Test live URL, then View tested page. The live test adds a Screenshot tab, and More info lists the resources Google loaded and any JavaScript console messages.

For a site you can't add to Search Console, Google's Rich Results Test fetches any public URL as Google-InspectionTool, with a smartphone user agent by default, and works from the rendered source. Both tools tell you exactly what Google sees. Neither tells you anything about GPTBot, which is why the curl check above still matters.

How to compare raw HTML vs rendered HTML

Save both versions and compare the parts crawlers use, not the whole files. Running diff on raw HTML is mostly useless, because minified markup often sits on one line and diff reports a single changed line the length of the page. This loop pulls out what matters from each file:

for f in raw.html rendered.html; do
  echo "== $f"
  grep -oiE '<title[^>]*>[^<]*' "$f" | head -1
  grep -oiE '<meta[^>]*robots[^>]*>' "$f"
  grep -oiE '<link[^>]*canonical[^>]*>' "$f"
  echo "links: $(grep -oiE '<a [^>]*href=' "$f" | wc -l)"
  echo "json-ld: $(grep -oi 'application/ld+json' "$f" | wc -l)"
  echo "words: $(perl -0777 -pe 's/<(script|style)\b.*?<\/\1>//gis; s/<[^>]+>/ /g' "$f" | wc -w)"
done

The word count is rough, since it still includes navigation and footers. The ratio is what tells you something. A raw count of 40 against a rendered count of 1,200 means crawlers that skip JavaScript get an empty page.

Check What a problem looks like Why it matters
Title The raw <title> is generic and scripts replace it Non-rendering crawlers keep the raw one
Meta robots noindex in the raw HTML that a script removes Google may skip rendering, so the page stays out of the index
Canonical Missing from the raw HTML, or changed by a script Google advises against changing the canonical with JavaScript
Main text Key paragraphs, prices or specs only in rendered.html AI crawlers can only quote what is in the raw file
Internal links Far fewer <a href> links in raw.html Crawlers find your other pages through links they can see
Structured data JSON-LD only in rendered.html Google can read injected JSON-LD, but non-rendering crawlers never get it

What to fix when the two differ

Put anything a crawler needs into the server response and leave only the interactive parts to the browser. That means the main content, real <a href> links, JSON-LD written into the template instead of a tag manager, and a title, canonical and robots value that no script touches. Our JavaScript SEO guide covers server-side rendering, static generation and hydration, and how to pick between them.

Don't chase every difference. A chat widget or a cookie banner that exists only after rendering costs you nothing.

Compare raw and rendered HTML on your own page

AI Crawler View does the whole comparison above in one run. It fetches your URL with a plain HTTP request that runs no JavaScript, then loads the same URL fresh in a headless browser and compares the two. It reports how much of the rendered page's main content is already in the raw HTML, and lists the headings, paragraphs, prices and structured data types that appear only after JavaScript runs.

It also flags a title, meta description or canonical that a script sets or changes, a noindex that a script adds or removes, internal links that only exist after rendering, and pages where a plain request gets an error or a bot check while a browser gets the page. A run costs 8 credits, and if the browser version can't be loaded the credits are not kept. It doesn't check robots.txt rules or prove what a specific bot fetched.

Keep reading