What is Googlebot and how does it crawl your site?
·8 min read
Googlebot is Google's web crawler, the program that downloads pages so Google Search can index and rank them. It runs as two crawlers under one name, Googlebot Smartphone and Googlebot Desktop, and the smartphone one makes most of the requests because Google indexes the mobile version of most sites. Googlebot finds URLs through links and sitemaps, fetches the HTML, renders it in a current version of Chromium and passes the result to indexing. A page Googlebot cannot fetch is missing from Google's results, and from AI Overviews and AI Mode as well.
Most Googlebot trouble is self-inflicted. A robots.txt line copied from a staging site, a firewall rule that rate-limits anything that is not a browser, a 503 maintenance page left up over a long weekend. Knowing what Googlebot asks for, and how it reacts to your server's answers, makes each of these easy to find.
What is Googlebot? Two crawlers, one name
Googlebot is the crawler for Google Search, and it comes in a phone version and a desktop version. Googlebot Smartphone simulates a visitor on a mobile device. Googlebot Desktop simulates one on a computer. Google's Googlebot documentation says Search primarily indexes the mobile version of most sites, so the majority of crawl requests come from the smartphone crawler and a minority from the desktop one.
That is mobile-first indexing, and its practical meaning is simple. Google ranks the page your phone visitors get. If your mobile template hides the comparison table, drops the reviews or trims the internal links that the desktop page has, Google works from the thinner version.
Both crawlers answer to the same robots.txt token, Googlebot, so you cannot allow one and block the other. Blocking Googlebot also reaches past the ten blue links. Google says it affects Discover, Google Images, Google Video and Google News too.
Googlebot user agent strings
Every Googlebot user agent contains Googlebot/2.1 and a link to google.com/bot.html. These are the current strings from Google's list of common crawlers, where W.X.Y.Z stands for the Chrome version Googlebot runs at the time of the request.
Googlebot Smartphone
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
Googlebot Desktop
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36The Chrome version moves whenever Google updates its renderer, so never match the full string. Match on Googlebot in log filters, firewall rules and analytics exclusions. And never trust the string on its own. Any script can send it, which is why verification works from the IP address instead.
How Googlebot finds, crawls and indexes pages
Googlebot discovers URLs from links on pages it already knows and from sitemaps. Google says there is no central registry of web pages, so a page with no internal links pointing to it and no sitemap entry may never be found.
Once a URL is queued, it moves through three steps.
- Crawl. Googlebot requests the URL, follows any redirects and downloads the HTML. CSS, JavaScript and images referenced by the page are separate fetches.
- Render. Pages that return 200 join a render queue, where a headless, evergreen Chromium runs the JavaScript. Google says a page can wait there a few seconds or longer, and a page with noindex may skip rendering entirely.
- Index. Google processes the rendered page, groups duplicates under one canonical URL and decides whether to store it.
Rendering is where JavaScript-heavy sites lose text and links. If your main content or navigation only appears after scripts run, start with our guide to JavaScript SEO. Google can also know about a URL and keep putting off the crawl. Discovered, currently not indexed covers that status and what moves it.
How much of a page Googlebot reads
Googlebot for Search reads the first 2 MB of an HTML file or other supported text file, and the first 64 MB of a PDF. At the cutoff it stops downloading and sends only what it already has to indexing. The limit counts uncompressed bytes, so gzip or Brotli on the wire buys you no extra room.
Older guides quote 15 MB. That number is still in Google's docs, but as the default for its crawlers in general, and the Googlebot page gives 2 MB for Search. Each CSS or JavaScript file the page references is fetched separately with its own limit, so external files are rarely the problem. Inline code is. A framework that inlines a large JSON state object for hydration, plus inline SVG icons and base64 images, can push footer links and JSON-LD past the line. Our page size checker measures a page against the 2 MB limit and shows what falls past it.
Google's crawlers use HTTP/1.1 or HTTP/2, whichever performs better for your site, with HTTP/1.1 as the default. HTTP/2 can save server resources, but Google says it gives no ranking boost. If HTTP/2 crawling causes trouble, your server can answer Google's HTTP/2 requests with a 421 status to opt out.
Googlebot crawl rate and crawl budget
Googlebot sets its crawl rate from how your server responds. For most sites, Google says Googlebot shouldn't visit more than once every few seconds on average. Fast, steady responses let the rate rise. Slow responses, 5xx errors and 429 rate-limit responses make it back off.
Those responses are also the only brake you control. Google removed the crawl rate limiter from Search Console on January 8, 2024, and Googlebot ignores crawl-delay in robots.txt. If Googlebot really is overloading a server, Google's advice is to answer its requests with 500, 503 or 429 for a short time, optionally with a Retry-After header. Keep it short. Google warns that serving those codes to the same URLs for more than a day or two can get them dropped from the index.
The accidental version is more common. A CDN or bot-protection rule that answers Googlebot with a 503 or 429 looks to Google exactly like a struggling server, so it crawls less. Visitors see a normal site, so nobody notices. The Crawl Stats report in Search Console, under Settings, breaks Googlebot's requests down by response code and is the quickest place to catch it.
Crawl budget, which Google defines as the set of URLs it can and wants to crawl on your site, matters only at scale. Google's crawl budget guide is written for sites with over a million unique pages that change weekly, sites with over 10,000 pages that change daily, and sites with a large share of URLs stuck in Discovered, currently not indexed. Google calls those numbers rough estimates. On a 300-page site, a "crawl budget problem" is nearly always thin pages or weak internal linking under another name. And Google does not accept payment to crawl a site more often.
Googlebot and AI Overviews
AI Overviews and AI Mode draw on pages Googlebot has crawled and Google has indexed. Google's AI features guide says a page must be indexed and eligible to show with a snippet to appear as a supporting link, and lists no extra technical requirements. Your robots.txt rules for Googlebot therefore decide AI Overview and AI Mode eligibility along with your rankings.
That makes blocking Googlebot the wrong tool for keeping content out of AI answers. The narrower controls are nosnippet, data-nosnippet and max-snippet, which limit what Google may quote, and the Search generative AI setting in Search Console. Since August 31, 2026 every site can use that setting to leave AI Overviews, AI Mode and Discover's generative AI features without any effect on ranking. We cover the trade-offs in how to turn off AI Overviews for your site.
Google-Extended is a separate control. It is a robots.txt token with no crawler or user agent of its own, and it decides whether content Google crawls may be used to train Gemini models and for grounding. Google says it does not affect inclusion or ranking in Search. Our Google-Extended guide has the exact rules.
Google crawlers that are not Googlebot
Google runs several other crawlers with similar names, and they show up in the same logs. These are the ones people most often confuse with Googlebot.
| Crawler | robots.txt token | What it does |
|---|---|---|
| Googlebot-Image and Googlebot-Video | Googlebot-Image, Googlebot-Video |
Fetches image and video files for Google's image and video features |
| GoogleOther | GoogleOther |
Generic crawler for Google product teams, including internal research |
| Google-InspectionTool | Google-InspectionTool |
Runs URL Inspection in Search Console and the Rich Results Test |
| AdsBot | AdsBot-Google |
Checks landing page quality for Google Ads |
| Google-Extended | Google-Extended |
Control token for Gemini training and grounding, with no crawler of its own |
Two of these catch people out. AdsBot ignores the User-agent: * group, so a blanket Disallow: / does not stop it. It follows only rules that name AdsBot-Google or AdsBot-Google-Mobile. Google-InspectionTool requests are live tests someone ran in Search Console or the Rich Results Test, and Google says the crawler does not affect Search, so blocking it only breaks those tests. For OpenAI, Anthropic, Perplexity and the other AI bots, see our list of AI crawler user agents.
How to verify Googlebot
To verify Googlebot, check the IP address, never the user agent. Google's verification steps are a reverse DNS lookup that should return a host on googlebot.com, google.com or googleusercontent.com, then a forward lookup on that host that must return the original IP. For firewall allowlists, Google's published IP range files, starting with common-crawlers.json for Googlebot, are easier to maintain, and our guide to log file analysis has a script for both methods.
Check whether Googlebot can index your page
Our Indexing & Canonical Checker tells you whether Googlebot may crawl and index a URL, and names the signal that blocks it. It reads the HTTP status and redirect chain, the robots.txt rule that matches the URL for Googlebot, meta robots and X-Robots-Tag directives, and the canonical tag. If the canonical points to another URL, it loads that target too and checks whether it is indexable. It also reports the snippet controls that decide whether Google Search, AI Overviews and AI Mode, and Bing and Copilot may quote the page. Like Googlebot, it reads only the first 2 MB of HTML.
Each problem comes with the exact tag, header or robots.txt line to change, and a run costs 8 credits. It cannot confirm that Google has indexed the URL. Search Console's URL Inspection does that.