JavaScript SEO: why AI crawlers miss client-rendered content
·7 min read
JavaScript SEO is the work of making sure crawlers can read pages whose content, links or metadata are built by scripts. Google runs JavaScript in a separate rendering step after it crawls a page, but the major AI crawlers from OpenAI, Anthropic and Perplexity read only the HTML your server sends and never run your scripts. Anything a page adds in the browser, like product copy fetched from an API or links that only work through click handlers, can rank in Google and still be invisible to ChatGPT, Claude and Perplexity. The fix is to put that content in the server HTML with server-side rendering, static generation or prerendering.
Client-rendered sites got away with this for years because Google renders them. A pricing page that stays empty until React mounts can still rank, and still give GPTBot nothing to quote.
Does Google render JavaScript?
Yes, but not right away. Google's JavaScript SEO basics describes three phases. Googlebot crawls the URL and parses the HTML. The page waits in a render queue, for a few seconds or longer, and Google renders it with an evergreen Chromium, a headless Chrome kept current. The rendered HTML goes on to indexing.
Some pages skip rendering. A page that returns a non-200 status such as a 404 may not be rendered, and neither may a page whose HTML carries noindex. Rendering is also not browsing. Google doesn't interact with the page, so content that waits for a scroll or a click never loads for Googlebot. For how Googlebot crawls in general, see what Googlebot is and how it works.
Which AI crawlers run JavaScript?
Almost none of them. The clearest public evidence is a study Vercel published with MERJ on December 17, 2024, based on crawler traffic across Vercel's network. It found that none of the major AI crawlers rendered JavaScript. OpenAI's and Anthropic's crawlers did download JavaScript files, 11.50% and 23.84% of their fetches, but never executed them.
| Crawler | Company | Runs JavaScript? |
|---|---|---|
| Googlebot | Yes, in a separate rendering step | |
| Applebot | Apple | Yes, with a browser-based crawler |
| GPTBot, OAI-SearchBot, ChatGPT-User | OpenAI | No |
| ClaudeBot | Anthropic | No |
| PerplexityBot | Perplexity | No |
| Meta-ExternalAgent | Meta | No |
| Bytespider | ByteDance | No |
| CCBot | Common Crawl | No |
Gemini is the exception among assistants. Per the study, it uses Googlebot's infrastructure and gets the same rendering. ChatGPT-User is the OpenAI agent that visits a page on a user's behalf, so even a live lookup in ChatGPT gets your raw HTML.
The study is almost two years old, and crawlers change without announcements. OpenAI's crawler documentation says nothing about rendering. So check your server logs too. A bot that runs your scripts requests the API endpoints only your client code calls, not just the .js files. For every AI user agent and what it does, see our list of AI crawlers.
Client-side rendering SEO: what crawlers miss
Client-side rendering means the server sends an HTML shell and JavaScript builds the page in the browser. This is what a crawler that skips the scripts receives from a typical single-page app:
<!doctype html>
<html lang="en">
<head>
<title>Acme</title>
<script type="module" src="/assets/index-4f3c2a1b.js"></script>
</head>
<body>
<div id="root"></div>
</body>
</html>The title survives. The headings, copy, prices and links to the rest of the site all live in that bundle, so to GPTBot the page is one word long.
Full single-page apps are the obvious case. The quieter one is a server-rendered site with holes in it:
- Reviews, FAQs or specs fetched from an API after the page loads.
- Prices and stock filled in by a script.
- JSON-LD injected by a tag manager. Google reads it after rendering. AI crawlers never see it.
- Menus and pagination that navigate through click handlers instead of links.
- A title or canonical that a client-side router sets after load.
Tabs and accordions are fine when their text is in the HTML and hidden with CSS. They are a problem when a click fetches the text.
How to test whether your content needs JavaScript
Fetch the page without a browser and search the response for a sentence from the main text. If curl can't find it, a crawler that doesn't run scripts can't either.
curl -sL -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
https://example.com/pricing | grep -c "Billed annually"A count of 0 means the sentence isn't in the HTML. Pick text from the middle of the page, because titles and footers are often in the shell even when the content isn't. On a 0, rerun without the grep. A 403 or a challenge page where your HTML should be means a bot rule caught the request, a separate problem from rendering. For other ways to put raw and rendered HTML side by side, follow how to view rendered source.
Server-side rendering SEO: put the content in the first response
Server-side rendering is the fix for most sites. The server runs your components and sends finished HTML, then JavaScript hydrates that HTML in the browser so the page stays interactive. Crawlers read the HTML. Visitors get the app.
| Approach | When the HTML is built | Use it for |
|---|---|---|
| Server-side rendering (SSR) | On every request | Search results, live stock, anything per visitor |
| Static generation (SSG) | At build time, optionally rebuilt on a timer | Marketing pages, docs, blog posts, most product pages |
| Prerendering | At build time, by saving the client app's output from a headless browser | A single-page app you can't move to a new framework yet |
Next.js, Nuxt, Astro, SvelteKit and Remix all render on the server, and most also build static pages. In Next.js the change is often moving one fetch from the browser to the server. This client component sends crawlers an empty list:
"use client";
import { useEffect, useState } from "react";
export default function Pricing() {
const [plans, setPlans] = useState([]);
useEffect(() => {
fetch("/api/plans").then((res) => res.json()).then(setPlans);
}, []);
return <ul>{plans.map((plan) => <li key={plan.id}>{plan.name} {plan.price}</li>)}</ul>;
}This server component puts the plans in the HTML and rebuilds the page at most once an hour:
// app/pricing/page.tsx
export const revalidate = 3600;
export default async function Pricing() {
const plans = await getPlans(); // runs on the server
return <ul>{plans.map((plan) => <li key={plan.id}>{plan.name} {plan.price}</li>)}</ul>;
}Skip dynamic rendering, the setup that detects bots by user agent and serves them a prerendered copy while people get the client app. Google's dynamic rendering page now calls it "a workaround and not a long-term solution" and recommends server-side rendering, static rendering or hydration instead. It also fails quietly for AI crawlers, since any bot missing from your user agent list gets the empty shell. Build-time prerendering avoids that because everyone gets the same HTML.
A Markdown copy for AI agents doesn't fix it either. Crawlers fetch your page URLs and read the HTML that comes back. Whether Markdown is worth adding on top is covered in Markdown vs HTML.
JavaScript SEO fixes for links, metadata and lazy-loading
Server rendering handles the content. These three details still break on server-rendered sites, usually because one component was written for the browser only.
Use real links, not click handlers
Google can only discover links that are <a> elements with an href attribute. A crawler that doesn't run scripts can't fire a click handler at all, so a link without an href leads nowhere.
<!-- Crawlers can follow this -->
<a href="/pricing">Pricing</a>
<!-- Crawlers cannot follow these -->
<span onclick="router.push('/pricing')">Pricing</span>
<a onclick="goTo('pricing')">Pricing</a>Framework link components such as Next.js <Link> render a real <a href>. Use them instead of calling router.push from a button.
Put title, canonical and robots meta in the server HTML
The title, meta description, canonical and robots meta tag belong in the HTML the server sends, with their final values.
<head>
<title>Pricing plans | Acme</title>
<meta name="description" content="Three plans, billed monthly or annually.">
<link rel="canonical" href="https://acme.com/pricing">
</head>Google's JavaScript guide is direct about two of these. Don't use JavaScript to change the canonical to a different URL. And don't count on JavaScript to remove a noindex. When Google finds noindex in the HTML it may skip rendering, so the script that removes the tag may never run and the page stays out of the index. If you want a page indexed, leave noindex out of the original HTML.
Adding noindex with JavaScript fails the other way round. Google sees the tag only after rendering, and crawlers that don't render never see it.
Lazy-load without hiding content
Google's lazy-loading guidance asks for content to load when it enters the viewport, through the browser's built-in lazy-loading or the IntersectionObserver API, and never to wait for a scroll or a click. It also says not to lazy-load what people see when the page opens.
That works for Google. AI crawlers set a higher bar, because text fetched by IntersectionObserver is still text fetched by a script. Keep text in the HTML and lazy-load images and embeds, which is where the bytes are anyway. Native lazy-loading keeps the image URL in the markup, so even a crawler that never renders sees it:
<!-- The URL is in the HTML. The browser delays the download. -->
<img src="/img/crawl-volume.png" alt="Monthly crawl volume by bot" loading="lazy" width="800" height="450">
<!-- The URL only exists after a script copies it into src. -->
<img data-src="/img/crawl-volume.png" class="lazyload" alt="Monthly crawl volume by bot">Infinite scroll needs real pages behind it. Give each chunk a stable URL such as /blog?page=2 and link to it with an <a href>, so crawlers can reach items past the first screen.
Check what AI crawlers get from your page
AI Crawler View runs this comparison on any public URL. It fetches the raw HTML the way a crawler that skips scripts gets it, renders the same page in a headless browser and compares the two. You get the share of main content present in the raw HTML, the headings, structured data types, prices and paragraphs that appear only after JavaScript, and how many internal links exist only after scripts run. It also flags a title, meta description, canonical or noindex that JavaScript sets, changes or removes, and pages where a plain request gets a bot check or an error while a browser gets the page.
Each run costs 8 credits, and if the rendering step fails the run is refunded. It doesn't check robots.txt rules or prove which bots fetched the page. Your server logs answer that.