GPTBot: what OpenAI's crawlers do and how to allow or block them

·7 min read

GPTBot is the crawler OpenAI uses to collect public pages that may be used to train its foundation models, and you control it with a User-agent: GPTBot group in robots.txt. It is one of three OpenAI agents that matter to site owners. OAI-SearchBot crawls pages so they can appear in ChatGPT search, and ChatGPT-User fetches a page when a person asks ChatGPT to open it. Blocking GPTBot keeps your future pages out of training and leaves ChatGPT search alone, which is what most people searching "block gptbot" actually want.

A fourth agent shows up in some logs. OAI-AdsBot visits only pages submitted as ads on ChatGPT, checks that they are safe, and according to OpenAI its data is not used to train foundation models. If you don't advertise on ChatGPT, you won't see it.

What GPTBot, OAI-SearchBot and ChatGPT-User each do

Each agent has one job and its own robots.txt token, and OpenAI's crawler documentation says each setting works independently of the others.

Agent What OpenAI says it does Follows robots.txt What blocking it costs IP list
GPTBot Crawls content that may be used to train generative AI foundation models Yes Nothing visible today. Future crawls are left out of training gptbot.json
OAI-SearchBot Surfaces websites in ChatGPT search Yes Your pages stop appearing in ChatGPT search answers, apart from navigational links searchbot.json
ChatGPT-User Fetches pages for actions users take in ChatGPT, Custom GPTs and GPT Actions Not reliably, because a person starts each request ChatGPT can't open your page when someone asks it to chatgpt-user.json
OAI-AdsBot Checks pages submitted as ads on ChatGPT Not stated Only matters if you run ChatGPT ads adsbot.json

Search opt-outs go through OAI-SearchBot only. OpenAI says ChatGPT-User plays no part in what appears in search, so neither a ChatGPT-User rule nor a GPTBot rule takes you out of ChatGPT search.

Should you block GPTBot?

Block GPTBot if your content is the thing you sell. If your pages exist to sell a product or a service, leave it allowed.

A GPTBot block tells OpenAI not to train future foundation models on your pages, and GPTBot stops fetching the paths you disallow. It leaves these alone.

  • Models already trained. A robots.txt rule applies to crawls made after it exists.
  • ChatGPT search, which runs on OAI-SearchBot, and the pages ChatGPT-User opens for people.
  • Other training crawlers. CCBot, which builds the public Common Crawl archive, answers to its own token, as our CCBot guide explains.
  • Your Google rankings.

The case for blocking is a bad exchange rate. Cloudflare's crawl data puts OpenAI at about 1,091 pages crawled for every visit it referred in July 2025, down from about 1,217 in January. That figure covers OpenAI's crawling as a whole, not GPTBot alone. A publisher of paid reporting, datasets or course material gets almost nothing back from training access.

The case against is about memory. When ChatGPT answers without searching, everything it says about your company comes from training data. Block GPTBot and newer models learn about you from what other sites say, not from your own pages. For a SaaS company or an agency, that is usually the worse outcome.

Make the opt-out in robots.txt, not only at the firewall. OpenAI says that when a site allows both GPTBot and OAI-SearchBot, it may use the results of one crawl for both purposes. If robots.txt still allows GPTBot and you only drop its requests at the CDN, OpenAI's own wording leaves training allowed through the search crawl.

How to block GPTBot in robots.txt

Add a group for the GPTBot token with Disallow: / to the robots.txt at the root of every host you want covered. Pick the version that matches your decision. Our guide to blocking AI crawlers covers the other vendors.

User-agent: GPTBot
Disallow: /

That is the whole change for most sites. OAI-SearchBot and ChatGPT-User have no group here, so they keep following your User-agent: * rules. If your * group disallows everything, as some files left over from staging do, give OAI-SearchBot its own group listing what it may crawl.

Keep GPTBot out of some folders only

User-agent: *
Disallow: /cart/
Disallow: /account/

User-agent: GPTBot
Disallow: /cart/
Disallow: /account/
Disallow: /research/
Disallow: /courses/

Repeat your * rules inside the GPTBot group. Once GPTBot has a group of its own it ignores the * group entirely, the precedence mistake our AI crawler list walks through.

Block GPTBot and OAI-SearchBot

User-agent: GPTBot
User-agent: OAI-SearchBot
Disallow: /

This removes you from ChatGPT search answers as well as training. OpenAI notes that sites which opt out of OAI-SearchBot can still appear as navigational links.

Stop ChatGPT-User

A robots.txt rule is the wrong tool here, because OpenAI says robots.txt may not apply to requests a user starts. Blocking ChatGPT-User takes a server or CDN rule on the user agent, such as this one in Nginx.

if ($http_user_agent ~* "chatgpt-user") {
    return 403;
}

Few sites should do this. Each blocked request is a person who asked ChatGPT to read your page and got an error instead.

OpenAI's crawler page doesn't mention Crawl-delay or Content-Signal lines, so don't count on either to change what its agents do. Write Disallow rules for the tokens above.

How long does a GPTBot robots.txt change take?

OpenAI says ChatGPT search takes about 24 hours to adjust after you update robots.txt. It gives no figure for GPTBot. To confirm a block took effect:

  1. Load your robots.txt in a private window and check the new rules are live, since a CDN can keep serving a cached copy.
  2. Do the same on every subdomain that serves pages. Each host reads only its own file.
  3. After two or three days, count GPTBot requests for anything other than /robots.txt with the log command below. The count should fall to zero.
  4. If disallowed pages are still being fetched, check the requesting IPs against gptbot.json before you blame OpenAI.

GPTBot keeps fetching robots.txt after a block. Rereading the file is how it knows to stay out.

What is the GPTBot user agent?

OpenAI publishes these strings and warns that the version numbers may change.

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot

When OAI-SearchBot fetches robots.txt it adds robots.txt; before the URL, and OpenAI says GPTBot may do the same. The marker separates robots.txt fetches from page crawls even in logs that don't record the path.

The format has changed before. The crawler-user-agents project recorded GPTBot/1.0 in August 2023 with "compatible" inside the parentheses. A filter written for that exact string misses current traffic, so match the token without case sensitivity and ignore the rest.

How to verify a request is really from OpenAI

Check the source IP against the list OpenAI publishes for that agent. Any scraper can send a GPTBot user agent, but only OpenAI sends it from these ranges.

Agent IP list IPv4 ranges on October 10, 2026
GPTBot openai.com/gptbot.json 18
OAI-SearchBot openai.com/searchbot.json 39
ChatGPT-User openai.com/chatgpt-user.json 234
OAI-AdsBot openai.com/adsbot.json 2

All four files listed only IPv4 ranges when we pulled them, and OpenAI's crawler page describes no reverse DNS check, so the IP lists are the test. Each file carries a creationTime, and the ChatGPT-User list was dated October 7, 2026, so refresh any copy you keep.

This script checks one IP against all four lists and names the agent it belongs to.

python3 - 74.7.227.12 <<'EOF'
import ipaddress, json, sys, urllib.request

ip = ipaddress.ip_address(sys.argv[1])
for bot in ["gptbot", "searchbot", "chatgpt-user", "adsbot"]:
    data = json.load(urllib.request.urlopen(f"https://openai.com/{bot}.json"))
    nets = [p.get("ipv4Prefix") or p.get("ipv6Prefix") for p in data["prefixes"]]
    if any(ip in ipaddress.ip_network(net) for net in nets):
        print(f"{ip} is in {bot}.json")
        break
else:
    print(f"{ip} is not in any OpenAI list")
EOF

Behind Cloudflare or a load balancer, the first field of your log may be the proxy's address. Record CF-Connecting-IP or the first X-Forwarded-For address, or every real OpenAI request will fail this check.

Use the lists to let traffic in, too. OpenAI recommends allowing OAI-SearchBot in robots.txt and allowing requests from its published ranges. If your CDN challenges unknown bots, a perfect robots.txt still leaves OAI-SearchBot facing a challenge page, so add searchbot.json to its allowlist.

How to find GPTBot, OAI-SearchBot and ChatGPT-User in your logs

Search the access log for each token and read the counts separately, because each agent's traffic means something different. A GPTBot request in a combined log looks like this.

74.7.227.12 - - [09/Oct/2026:14:02:51 +0000] "GET /blog/pricing-guide HTTP/1.1" 200 24511 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"
# Requests per OpenAI agent and version
zgrep -hoiE "(gptbot|oai-searchbot|chatgpt-user|oai-adsbot)/[0-9.]+" /var/log/nginx/access.log* | tr 'A-Z' 'a-z' | sort | uniq -c | sort -rn

# Pages ChatGPT-User opened for people
zgrep -hi "chatgpt-user" /var/log/nginx/access.log* | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

# Status codes OAI-SearchBot received
zgrep -hi "oai-searchbot" /var/log/nginx/access.log* | awk '{print $9}' | sort | uniq -c | sort -rn

# GPTBot requests other than robots.txt, per day
zgrep -hi "gptbot" /var/log/nginx/access.log* | awk '$7 != "/robots.txt" {print substr($4, 2, 11)}' | sort | uniq -c

# IPs fetching robots.txt with OpenAI's marker
zgrep -hi "; robots.txt;" /var/log/nginx/access.log* | awk '{print $1}' | sort | uniq -c

Read the results by agent. GPTBot volume shows what OpenAI collects for training and says nothing about what users see. OAI-SearchBot is your ChatGPT search index, so any 403 or 429 in its status codes is the first thing to fix. ChatGPT-User is the closest thing to demand data you get from OpenAI, since each hit started with a person in ChatGPT or a custom GPT asking for that page.

Expect GPTBot near the top of your AI bot traffic. In Cloudflare's data its share of AI-only crawler traffic rose from 11.9% in July 2024 to 28.1% in July 2025, the largest of any AI crawler. To skip the shell work, paste a log into our AI Bot Log Analyzer, which separates OpenAI's agents, checks each IP against OpenAI's lists and flags impostors.

Check your robots.txt for OpenAI's crawlers

Our robots.txt checker fetches your live robots.txt and reports whether GPTBot, OAI-SearchBot and ChatGPT-User may crawl the URL you enter, along with 11 other AI tokens such as ClaudeBot, PerplexityBot and Google-Extended. For each one it names the rule that decided the result and says when the crawler fell back to your * group. A blocked search or user crawler comes back as a warning. A blocked training crawler like GPTBot comes back as a note, since that is usually deliberate.

Give it a page URL and a user agent such as GPTBot and it also tests that exact path, the quick way to confirm a folder-level rule. It reads Content-Signal lines and flags duplicate user-agent groups, Allow and Disallow rules that tie, a missing * group and HTML served in place of robots.txt. It does not pretend to be GPTBot, so it shows what your file says, not what OpenAI's crawler does. A run costs 10 credits.

Keep reading