Google-Extended: what it blocks and what it doesn't

·7 min read

Google-Extended is a robots.txt token that tells Google whether content it crawls from your site may train future Gemini models and ground answers in Gemini Apps and in Grounding with Google Search on Vertex AI. Google-Extended is not a crawler, so it never shows up in your logs. Blocking it doesn't remove your pages from Google Search, AI Overviews or AI Mode, and Google says it isn't a ranking signal. To stay out of AI Overviews and AI Mode, use the Search generative AI setting in Search Console or a nosnippet rule.

The name invites two wrong guesses. People expect a bot they can block at the firewall like GPTBot, and they expect it to cover everything Google does with AI. It does neither, and the second guess costs more. A site that blocks Google-Extended to get out of AI Overviews stays in AI Overviews and leaves the Gemini app's answers instead.

What is Google-Extended?

Google-Extended is a product token, a name you can put on a User-agent line, with no crawler behind it. Google introduced it on September 28, 2023 for Bard and the Vertex AI generative APIs, according to Google's announcement. Bard became Gemini, and the token now covers Gemini.

Google's common crawlers documentation says it has no user agent string of its own. Google's existing crawlers do the fetching, and the token is used "in a control capacity". A crawler such as Googlebot fetches the page, then Google checks your Google-Extended rules to decide what else it may do with that copy.

So no request ever carries the name. You can't confirm the rule from access logs or block it with a firewall rule, and disallowing it doesn't reduce Google's requests to your server.

What Google-Extended controls

A Disallow for Google-Extended covers two uses of your content.

Training. Google won't use the content to train future generations of the Gemini models behind Gemini Apps and the Vertex AI API for Gemini. Note the word future. Models trained before you added the rule keep what they learned.

Grounding. Google defines grounding as handing content from the Google Search index to the model at prompt time, to make its answer more accurate. With Google-Extended blocked, Google stops doing that with your pages in Gemini Apps and in Grounding with Google Search, the Vertex AI feature developers use to connect their own Gemini apps to Search results.

The grounding half is easy to miss. A training opt-out sounds like it costs nothing, but this one also pulls your pages out of the answers the Gemini app builds from search results. If buyers ask Gemini for recommendations in your category, that is a real loss. Our guide to Gemini SEO covers how the app picks the sites it cites.

Google-Extended does not block AI Overviews

AI Overviews and AI Mode are part of Google Search, and Google says Google-Extended has no effect on a site's inclusion in Search. Google's AI features documentation names Googlebot's robots.txt rules as the access control for those features. Disallowing Google-Extended leaves your pages exactly as eligible as before.

Two other controls do keep you out. The Search generative AI setting, under Settings in Search Console, reached every site on August 31, 2026 and defaults to Include. Set to Exclude, it keeps the property out of AI Overviews, AI Mode and Discover's generative AI features without changing ranking. For a single page, nosnippet stops Google from using the content as a direct input for AI Overviews and AI Mode, per the robots meta tag documentation, and costs the page its text snippet in regular results. Our guide to turning off Google AI Overviews walks through both.

The help page for the Search setting says it doesn't affect AI training and points to Google-Extended for that. Getting out of both takes both switches.

How to block Google AI training with robots.txt

Add a group for the token to the robots.txt at the root of each host. Google matches user agent names without regard to case.

User-agent: Google-Extended
Disallow: /

To opt out one section, disallow only its paths.

User-agent: Google-Extended
Disallow: /research/
Disallow: /members/

Google generally caches robots.txt for up to 24 hours, so allow a day for the change. A robots.txt on a subdomain covers only that subdomain, so blog.example.com needs its own group.

Mistakes that block Search or miss the token

The mistake that costs the most is putting Google-Extended in the same group as Googlebot. Consecutive User-agent lines share one set of rules, so this file takes the site out of Google Search too.

# Wrong: blocks Google Search as well as Gemini
User-agent: Googlebot
User-agent: Google-Extended
Disallow: /

The reverse catches sites that run an allowlist, with Disallow: / under User-agent: * and open groups for Googlebot and Bingbot. Google's crawlers follow the most specific group that names them and use the * group only when none does, per Google's robots.txt rules. Google-Extended has no group in that file, so it inherits the * block and the site has quietly opted out of Gemini grounding. If you want Gemini to cite you, name the token.

User-agent: *
Disallow: /

User-agent: Googlebot
User-agent: Google-Extended
Allow: /

Google-Extended vs GoogleOther

GoogleOther is a real crawler, and blocking it is not how you opt out of Gemini training. Google calls it the generic crawler its product teams use to fetch public content, for example in one-off crawls for internal research and development, and says rules addressed to it don't affect any specific product. Its user agent looks like this.

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; GoogleOther) Chrome/W.X.Y.Z Safari/537.36

Google-Extended's definition covers content Google crawls from your site without naming a crawler, so it is the training control whichever Google bot made the fetch. Block GoogleOther if its traffic bothers you. Search won't notice.

Several other Google agents get confused with these two.

Name What it is What a Disallow does In your logs
Googlebot The Search crawler Google can't crawl the page for Search, AI Overviews or AI Mode Yes
Google-Extended A control token Opts out of Gemini training and grounding Never
GoogleOther Generic crawler for Google's product teams Affects no specific product, per Google Yes
Google-CloudVertexBot Crawler for Vertex AI Agents Stops crawls a site owner requested for an agent Only after such a request
Google-Agent Fetcher for agents acting on a user's request Usually nothing Yes
Google-GeminiNotebook Fetcher for sources a Gemini Notebook user adds Usually nothing Yes

Google-CloudVertexBot crawls only when a site's owner asks Google to crawl it for a Vertex AI agent, and Google says it has no effect on Search or other products. If it shows up and you never built an agent, ask whoever runs your Google Cloud account.

Google-Agent and the Gemini Notebook fetcher, whose old Google-NotebookLM user agent Google supported until August 2026, are on Google's user-triggered fetchers list. Google says these generally ignore robots.txt because a person asked for the page, so no group stops them. For the bots other companies run, see our AI crawler list.

Which Google control does what

Each control does one job and has its own side effect.

What you want Control What else changes
Keep content out of Gemini training Disallow Google-Extended Gemini Apps and Vertex AI stop grounding answers in it. Search is unchanged
Stay out of AI Overviews and AI Mode Search generative AI set to Exclude Discover's generative AI features too. Ranking and training unchanged, per Google
Keep one page's text out of AI Overviews nosnippet in a meta tag or X-Robots-Tag The page loses its text snippet in regular results
Leave Google Search entirely noindex Every click from Google Search

Google-Extended and the Search setting are independent, so each combination does something different. The one to avoid is Google-Extended blocked by someone who wanted out of AI Overviews. They stay in AI Overviews and leave the Gemini app's answers.

Should you block Google-Extended?

Leave it open if you want Gemini to recommend you. Software companies, agencies, local services and most stores gain more from a Gemini citation than they lose to training, and blocking gives up both.

Block it when the content is what you sell. Publishers, research firms and anyone licensing data have a reason to refuse training, and Google-Extended is the only Google control that does that while keeping you in Search. Know its limits going in. AI Overviews can still summarize the same pages unless you also set Exclude, and models already trained keep what they learned.

If you're refusing AI training across the board, Google-Extended is one group in a longer file. Our guide on how to block AI crawlers covers the rest.

Check your Google-Extended rules

Our robots.txt checker fetches your live robots.txt and shows whether 14 AI tokens, Google-Extended among them, may fetch the URL you enter. For each one it names the rule that decided it and whether the token fell back to your * group, which is how the allowlist mistake above shows up. The path test takes any crawler name and path, so you can test Googlebot, GoogleOther or Google-CloudVertexBot on the section you care about.

It also flags syntax errors, groups that disallow the whole site, duplicate groups and tied Allow and Disallow rules, and it reads any Content-Signal lines. A run costs 10 credits. It reads your rules only, so it can't see your Search Console setting or prove what Google does with your content.

Keep reading