What is a web crawler and how does it read your site?
·4 min read
A web crawler is a program that downloads web pages automatically, reads the links in them and queues those links to download next. Search engines and AI companies run website crawlers to build the index or dataset their answers come from. A crawler finds your pages through links and sitemaps, checks robots.txt before fetching, and reads either your raw HTML or the page after JavaScript runs, depending on the bot. A page it never finds a path to does not exist for that engine.
How a web crawler discovers URLs
A crawler discovers URLs from links on pages it already has and from XML sitemaps. Google's How Search works guide says there is no central registry of all web pages, so Google has to keep looking for new ones. It names two routes in, following a link from a known page or reading a sitemap.
Links do most of the work. The crawler starts from URLs it knows, often your homepage, pulls every link out of the HTML and queues the new ones. Each page it fetches feeds more URLs back in.
Only real links count. Google's link guidelines say it can reliably extract URLs only from <a> elements with an href attribute. A <div> with a click handler or a button that calls router.push() is invisible to the queue.
Sitemaps are the backup. They list URLs you want crawled, but Google's sitemap docs say listing a URL doesn't guarantee it gets crawled or indexed. A sitemap tells the crawler a page exists. A link tells it the page matters.
How a crawler fetches a page and obeys robots.txt
Before a well-behaved crawler fetches any URL on your host, it requests /robots.txt and checks whether its user agent may crawl that path. The rules are standardized in RFC 9309, published in September 2022. A few details from it cause real trouble:
- If robots.txt returns a 4xx error, the crawler may crawl anything.
- If it returns a 5xx error, the crawler must treat the whole site as disallowed.
- Crawlers may cache robots.txt, and should not use a cached copy for more than 24 hours.
- The rules "are not a form of access authorization." Bad bots ignore them.
A robots.txt that throws 500s after a CDN or firewall change stops compliant crawlers cold. A missing file lets everyone in.
Robots.txt controls crawling, not indexing. Google's robots.txt introduction says it is not a way to keep a page out of Google, and a blocked URL can still be indexed if other pages link to it. Use noindex for that. Our post on robots.txt Disallow covers the syntax.
Once allowed, the crawler requests the URL, follows redirects and downloads the response. A 200 moves on to processing, a 404 or 410 drops the page.
Do crawlers render JavaScript?
Googlebot does, the rest mostly don't. Googlebot queues pages that return 200 for rendering in a headless Chromium, then reads the rendered HTML, links included. Our post on what Googlebot is walks through that pipeline.
AI crawlers are different. A study Vercel published with MERJ in December 2024 found that none of the major AI crawlers it measured executed JavaScript. GPTBot, ClaudeBot and PerplexityBot read the raw HTML your server sends. If your navigation or body text appears only after scripts run, those bots see neither the content nor the links to your other pages. A client-rendered menu hides every page it links to, leaving the bot only your sitemap.
Search crawlers vs AI crawlers
Search crawlers fetch pages to build a ranked index. AI crawlers fetch pages to train models or to answer a user's question live. Both discover and fetch the same way, but each bot has its own user agent and robots.txt token.
| Search crawlers | AI crawlers | |
|---|---|---|
| Examples | Googlebot, Bingbot | GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot |
| Purpose | Index pages for search results | Model training, AI search, live lookups |
| Runs JavaScript | Googlebot yes | Mostly no |
| Controlled by | robots.txt token per bot | robots.txt token per bot |
Allowing Googlebot does not allow GPTBot, and blocking one does not block the other. For every AI bot, what it does and the token to use, see our list of AI crawlers.
Why crawlers miss pages with no internal links
A page with no internal links pointing at it is an orphan page, and crawlers miss it because nothing puts it in their queue. If it is in your sitemap, Google may still find it, though with no link signal telling it the page matters. If it is missing from the sitemap too, no compliant crawler will ever request it.
Orphans usually come from redesigns that drop a navigation level and campaign pages nobody linked. Pages four or more clicks deep have a milder version of the problem. Our guide to orphan pages covers how to fix each kind.
Find the pages a crawler cannot reach
The Orphan Page Finder crawls your site the way a search crawler does. It starts from the URL you enter, follows same-host links breadth-first up to 150 pages, skips URLs robots.txt blocks for Googlebot and doesn't expand pages marked nofollow. In parallel it reads the sitemaps your robots.txt declares, or the default sitemap paths, and compares the two lists.
It returns sitemap pages nothing links to, indexable pages missing from the sitemap, pages four or more clicks deep, broken internal links and sitemap URLs that redirect, fail or carry noindex. It does not run JavaScript, so links added after load are not followed, and it does not see links from other sites. On a site larger than 150 pages it covers the pages it reached first. Each run costs 20 credits.