Page size and SEO: how big is too big for Googlebot?
·4 min read
Page size matters for SEO at one hard line. Googlebot for Search reads the first 2 MB of an HTML file, counted uncompressed, and indexes only that part, so anything after the cutoff does not exist to Google. Almost no hand-written page gets near it. The pages that do cross it are carrying megabytes of inline hydration data, base64 images or inline SVG, and the fix is to move those bytes out of the HTML.
Googlebot's page size limit: 2 MB, not 15 MB
Googlebot fetches the first 2 MB of a supported file type and the first 64 MB of a PDF, according to Google's Googlebot documentation. When it reaches the limit it stops downloading and sends what it already has to indexing.
If you remember the Googlebot 15MB limit, you are not wrong about the number, only about who it applies to. Google's crawler overview says its crawlers and fetchers read the first 15 MB of a file by default, and that individual projects can set their own limits. Googlebot for Search sets 2 MB.
Three details decide whether you are at risk:
- The limit applies to uncompressed bytes. Gzip or Brotli shrinks the transfer but gives you no extra room.
- Each CSS file, JavaScript file and image the page references is fetched separately with its own limit. A 3 MB JavaScript bundle loaded by
<script src>does not count against your HTML. - Order matters. Footer links, JSON-LD at the end of
<body>and the last third of a long article are the first things to fall off.
More on fetching and rendering in what Googlebot is and how it crawls.
What actually pushes HTML size past 2 MB
Text alone rarely approaches 2 MB, because 2 MB of plain prose is a few hundred thousand words. The HTML size problem comes from data embedded in the document.
Hydration payloads. Frameworks that render on the server often inline the page's state as JSON so the client can hydrate. Next.js writes it into self.__next_f scripts or __NEXT_DATA__, Nuxt into window.__NUXT__. A product listing that serializes every variant, review and related item can put a megabyte of JSON into the HTML.
Base64 images. A data:image/png;base64,... URI is about a third larger than the file it encodes, and it lives inside the HTML. A handful of large inlined photos can cross the line by themselves. Browsers cannot cache them separately, and Google does not index them as images.
Inline SVG. An icon set pasted into every card, or a map drawn as an <svg> with thousands of path points, multiplies on listing pages that repeat it.
Inline CSS from page builders stacks on top, though it is rarely the main cause.
The sign to look for is HTML that is mostly code. A page whose readable text is 30 KB inside a 2.4 MB document has a data problem, not a content problem.
Page weight vs page speed
Page weight is the total bytes a page downloads, HTML plus every image, script, font and stylesheet. The 2 MB limit only cares about the HTML. Speed cares about everything, and the two fail in different ways.
A heavy HTML document delays everything that follows. The browser cannot discover the hero image or the stylesheet until it has parsed the markup that references them, and a big inline JSON blob has to be downloaded and parsed before hydration finishes. That shows up as slower Largest Contentful Paint (LCP) and worse interaction responsiveness.
Total page weight is a weaker signal than people expect. Waiting on the server and on resource discovery usually costs more LCP time than downloading the image, which is why our guide to improving LCP starts with TTFB and load delay rather than compression.
How to reduce page size
Measure first, then work from the largest block down.
- Move data out of the HTML. Send only what the first screen needs in the hydration payload and fetch the rest from an API after load.
- Replace base64 images with files. Serve them with
<img src>so the browser caches them and Google can index them. Keepdata:URIs for tiny placeholders. - Use an SVG sprite or image files for icons. Reference repeated icons with
<use href>or<img>instead of pasting the same paths into every card. - Serve CSS and JavaScript as external files. They are fetched separately and cached across pages.
- Paginate long lists. An archive page with 2,000 items belongs on 40 pages, each linked plainly.
- Compress the response. This does not help with Googlebot's limit, but it cuts transfer time for users.
- Lazy-load images below the fold. Never lazy-load the LCP image. A web.dev study by Felix Arntz and Rick Viscomi found that lazy-loading above-the-fold images delayed LCP.
<img src="/img/gallery-12.webp" alt="Oak desk, side view" width="800" height="600" loading="lazy">Keep the text itself in the HTML. Moving body copy behind a script to save bytes trades a size problem for a rendering one, which the JavaScript SEO guide covers.
Check your page size against Googlebot's limit
Our page size checker fetches a URL and measures its HTML in bytes against Googlebot's 2 MB limit. When the page is over, it counts the internal links, JSON-LD blocks and headings past the cutoff and the share of main text Googlebot misses. It breaks the HTML down into inline scripts, JSON-LD, inline styles, base64 images, inline SVG and text, lists the five largest inline scripts, and counts DOM elements and nesting depth against Lighthouse's thresholds.
It reads the HTML as served. It does not run JavaScript, measure images, scripts or styles loaded as separate files, or report Core Web Vitals. Each run costs 5 credits.