Can ChatGPT search the internet? When AI reads your site live
·7 min read
Yes, ChatGPT can search the internet, on every plan and for people who are not signed in. It decides by itself when a question needs current information, and you can force a search by choosing Search from the tools menu or typing / in the message box. When it skips the search, it answers from training data that ends at a knowledge cutoff, and that answer carries no links.
The split matters if you run a website. An answer from training data reflects whatever OpenAI collected months or years ago. An answer from a live search reflects your page as OpenAI's crawlers fetched it recently, and those crawlers read the HTML your server sends without running JavaScript. If your content appears only after scripts run, ChatGPT gets an empty page and has nothing of yours to quote.
Can ChatGPT search the internet by itself?
Yes. By default ChatGPT chooses when to search, and OpenAI's help center says it searches automatically when a question would benefit from current information from the web.
You can also make it search. Open the tools menu in the message box, choose View all tools, then Search. Typing / in the message box and picking Search does the same thing.
When it searches, ChatGPT rewrites your question into one or more shorter queries and sends them to third-party search providers. It reads what comes back and writes an answer with citations you can click. The query rewriting and source picking get their own walkthrough in how ChatGPT search works.
One exception to "every plan". Admins of Enterprise and Edu workspaces can switch web search off, and then ChatGPT will not search even when someone selects Search by hand.
Does ChatGPT have internet access without search?
No. The model itself has no connection to the internet. ChatGPT reaches the web through tools, and search is the main one. Take the tools away and it knows only what was in its training data.
This is why ChatGPT can give a confident answer that is out of date. Without a search it has no way to know your pricing changed in March or that you discontinued a product. It answers from what it learned, and the only sign the answer may be stale is the missing citations.
What is ChatGPT's knowledge cutoff?
ChatGPT's knowledge cutoff is the date its training data ends, and it depends on which model is answering. OpenAI lists the cutoff on each model's page in its API docs. GPT-5 lists September 30, 2024, and GPT-5.5 lists December 1, 2025. OpenAI ships new models several times a year, so there is no single date worth memorizing.
What an AI knowledge cutoff means for your site
Every language model has an AI knowledge cutoff, Claude and Gemini included. The cutoff is also a ceiling, not a guarantee. A page published a few weeks before it may be thin in the training data, and a small site may not be there at all. For anything recent or niche, a live search is the only way your current page gets into the answer.
Training data vs live search: how to tell which answer you got
Look for citations. A search answer shows its sources as links you can open, and an answer from training data has none.
| Answer from training data | Answer from a live search | |
|---|---|---|
| Where the text comes from | What the model learned during training | Pages fetched for this question |
| How current it is | Stops at the model's knowledge cutoff | As current as the page it fetched |
| Links to sources | None | Citations you can open |
| OpenAI crawler involved | GPTBot, long before your question | OAI-SearchBot for the index, ChatGPT-User on request |
| Your robots.txt control | Disallow GPTBot to stay out of future training | Allow OAI-SearchBot to appear in search answers |
| How fast a change shows up | When a new model is trained | About 24 hours after a robots.txt change, per OpenAI |
For a site owner, the right column is the one worth working on. You cannot edit what a model already learned. You can change what its crawlers find on your page tomorrow.
Which bots fetch your site for ChatGPT search?
Two OpenAI agents fetch pages for search answers, and a third collects training data. OpenAI's crawler docs describe all three.
OAI-SearchBot. It crawls pages for the index ChatGPT search draws from. Sites that block it will not be shown in ChatGPT search answers, though they can still appear as plain navigational links. OpenAI says a robots.txt change takes about 24 hours to reach search.
ChatGPT-User. It visits a page when a user's question leads ChatGPT there, or when a custom GPT calls an external action. It does not crawl on its own schedule, and because a person started the request, OpenAI says robots.txt rules may not apply.
GPTBot. It collects pages that may be used to train OpenAI's models. OpenAI says each setting is independent, so you can block GPTBot and stay in search.
This is how they identify themselves in your server logs. OpenAI notes the version numbers can change.
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbotIf you want to appear in ChatGPT search but keep your pages out of training, this is the robots.txt block:
User-agent: OAI-SearchBot
Allow: /
User-agent: GPTBot
Disallow: /A broad User-agent: * disallow or a CDN setting that blocks AI bots can shut OAI-SearchBot out without anyone noticing. You can test your robots.txt against OAI-SearchBot and the other AI crawlers to see which rule matches each bot.
Do AI crawlers run JavaScript?
No. In December 2024, Vercel and MERJ analyzed AI crawler traffic on Vercel's network and found that none of the major AI crawlers render JavaScript. The crawlers they checked included OAI-SearchBot, ChatGPT-User and GPTBot. ChatGPT's crawlers did download JavaScript files, 11.5% of their requests, but never ran them.
So a page that builds its content in the browser reads as empty. This is roughly what a client-rendered React app sends to OAI-SearchBot:
<body>
<div id="root"></div>
<script type="module" src="/assets/index-4f3c2a1b.js"></script>
</body>Everything a visitor sees comes from that script. The crawler gets an empty div, so ChatGPT either cites another page or describes you from whatever it learned in training.
Whole single-page apps are the obvious case. The partial cases are harder to spot. Prices fetched from an API after load, reviews from a third-party widget, FAQ answers inside a script-rendered accordion and JSON-LD injected by a tag manager all look fine in your browser. They are also the facts a buyer is most likely to ask ChatGPT about.
A quick test is to fetch the raw HTML and search it for a sentence you can see on the page:
curl -s https://example.com/pricing | grep -c "per seat"A count of 0 for text that is visible in your browser means crawlers that don't run scripts can't see it. The fix is server-side rendering or static generation, so the content is in the HTML response. Once it is, getting picked as the cited source is a separate job, covered in how to get cited by ChatGPT.
Google is the exception. The same study notes that Gemini uses Googlebot's infrastructure, which renders JavaScript.
Which other AI assistants search the internet live?
Claude, Perplexity and Gemini all search the web live, and each sends its own fetchers. Claude and Perplexity follow OpenAI's split, with one bot for the search index and one for user requests.
| Assistant | Crawler for the search index | Fetcher for user requests | Ran JavaScript in the Vercel and MERJ study |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | ChatGPT-User | No |
| Claude | Claude-SearchBot | Claude-User | No (ClaudeBot was tested) |
| Perplexity | PerplexityBot | Perplexity-User | No (PerplexityBot was tested) |
The robots.txt rules differ. Anthropic says blocking Claude-User stops Claude from retrieving your page when a user asks, while Perplexity says Perplexity-User generally ignores robots.txt because a person requested the fetch. Anthropic's third bot, ClaudeBot, collects training data, and our ClaudeBot guide covers how the three differ.
Gemini works differently. It grounds answers in Google Search's index, which Googlebot crawls. Google's Google-Extended token controls whether Gemini can use your content for training and grounding, and it has no effect on your inclusion or ranking in Google Search.
Check what AI crawlers see on your page
AI Crawler View fetches your page twice. Once as plain HTML, the way crawlers that don't run scripts get it, and once in a headless browser that runs the scripts. It compares the two and reports what share of the main content is in the raw HTML. It also lists the headings, prices, structured data and internal links that appear only after JavaScript, and flags a title, description, canonical or noindex that scripts change.
It catches the bot-protection case too, where a plain request gets an error or a challenge page while a browser gets the real page. Each finding comes with the fix, which is usually to move that piece into the server HTML.
AI Crawler View checks one URL per run and costs 8 credits. If the browser render fails, the run is refunded. It does not test robots.txt, so use the robots checker for that.