CCBot and Common Crawl: should you block it?

·7 min read

CCBot is the crawler behind Common Crawl, a nonprofit that publishes a snapshot of the web about once a month for anyone to download, and that archive is one of the most common sources of AI training data. Block CCBot if the text on your pages is your product, such as paid articles or courses. Leave it allowed if your pages exist to sell something else, because ChatGPT, Claude, Perplexity and Google find the pages they cite with their own crawlers, so a CCBot block costs you no AI search visibility. Either way, a block only stops future crawls, and pages already in published snapshots stay there.

What is CCBot?

CCBot is the web crawler that collects pages for Common Crawl's archive. Common Crawl is a 501(c)(3) nonprofit founded in 2007, and its home page puts the archive at over 300 billion pages, cited in over 10,000 research papers. The files sit in Amazon's public datasets program, where anyone can download them.

A new crawl comes out roughly every month. Common Crawl released 12 in 2025 and 9 in 2026 through September, according to its collection list, and each one takes about two weeks. The August 2026 crawl ran from August 7 to 20 and captured 2.14 billion pages from 33.1 million registered domains. About 0.7 billion of those URLs had never been crawled before.

It does not copy whole sites. The Common Crawl FAQ calls the dataset a sample of the web, built from a randomly selected subset of each site. Expect some of your URLs in any given crawl, and not always the same ones.

Common Crawl as AI training data

Common Crawl is one of the most widely used sources of AI training data. Mozilla Foundation's February 2024 report, Training Data for the Price of a Sandwich, looked at 47 large language models published between 2019 and October 2023 and found that at least 64% were trained on Common Crawl data. For GPT-3, over 80% of training tokens came from it.

Model builders rarely train on the raw archive. They train on cleaned datasets built from it, such as Hugging Face's FineWeb, which holds more than 18.5 trillion tokens of English text drawn from every Common Crawl dump since 2013. Common Crawl trains no models itself. Your pages reach model builders through copies of copies, which is why the block decision below is less clean than it looks.

What is the CCBot user agent?

CCBot identifies itself with this string, according to Common Crawl's FAQ.

CCBot/2.0 (https://commoncrawl.org/faq/)

Older logs show CCBot/1.0 (+https://commoncrawl.org/bot.html), and Common Crawl says it may raise the version number again, so match on the token CCBot. The robots.txt token is the same word.

Other bots borrow the name. Common Crawl's CCBot page warns about crawlers that falsely identify as CCBot. The real one runs from the ranges in ccbot.json, which in its August 11, 2026 version holds four IPv4 blocks totaling 28 addresses plus one IPv6 range. Its IPv4 addresses reverse-resolve to hostnames ending in crawl.commoncrawl.org. A request claiming to be CCBot from anywhere else is fake, and our list of AI crawlers walks through the reverse and forward DNS check.

Does CCBot respect robots.txt and Crawl-delay?

Yes, according to Common Crawl's FAQ. CCBot reads robots.txt before it fetches a page, obeys Crawl-delay, honors nofollow on links in your pages and picks up any sitemap listed in robots.txt. It rechecks the file periodically, so a new rule takes effect on its next check, not the moment you save.

Without a Crawl-delay it waits a few seconds between requests to the same site, and it slows down further when your server answers with a 429 or any 5xx status. It follows up to four redirects, or five for robots.txt. It does not run JavaScript or send cookies, so it archives the HTML your server returns and nothing more.

That matters for publishers. Common Crawl says it does not scrape paywalled material, but a paywall that hides article text with a script after the page loads hides nothing from a crawler that never runs the script. If the full article is in the HTML, CCBot gets the full article.

Should you block CCBot?

Block CCBot if you sell the words on your pages. Allow it if the pages exist to sell a product or service.

Block CCBot Allow CCBot
Citations in ChatGPT, Claude, Perplexity, AI Overviews No change No change
Future Common Crawl snapshots Your pages left out A sample of your pages each month
Pages in snapshots already published Stay there Stay there
New datasets and models built from future crawls Do not see your new pages Can include them
Load on your server None from CCBot A few seconds between requests, adjustable with Crawl-delay

The first row is the one people get wrong. ChatGPT search fetches with OAI-SearchBot, Claude with Claude-SearchBot, Perplexity with PerplexityBot and Google's AI features with Googlebot. A CCBot rule touches none of them.

The cost of blocking is slow and hard to see. Models trained on future Common Crawl derivatives will not have your new pages, so they learn less about your brand from your own words. Researchers lose your pages too, and the FAQ lists translation software, trend prediction and tracking disease spread among the archive's uses.

Blocking CCBot alone also buys little privacy. OpenAI, Anthropic, Meta and others run their own training crawlers, so if you want out of training sets, CCBot is one line among several. Our guide on how to block AI crawlers covers the rest in one file.

So for a SaaS company, an agency, a store or a local business, leave CCBot allowed. Your pages are marketing, and the open corpora that models learn from are cheap exposure. For a news publisher or anyone who licenses their text, block it along with the other training bots.

How to block CCBot in robots.txt

Add a CCBot group with Disallow: /, the rule Common Crawl's FAQ gives.

User-agent: CCBot
Disallow: /

To keep CCBot out of part of the site only, list those paths instead.

User-agent: CCBot
Disallow: /premium/
Disallow: /members/

If bandwidth is the problem, slow CCBot down instead of blocking it. Common Crawl's example value is 2, which means one request every 2 seconds.

User-agent: CCBot
Crawl-delay: 2
Disallow: /cart/
Disallow: /search

Copy your User-agent: * rules into any CCBot group you add. Under RFC 9309, a crawler obeys only the group that names it most specifically, so once CCBot has its own group it ignores everything under *. Each subdomain reads its own robots.txt, so add the rules on every host you want covered.

If something calling itself CCBot keeps crawling after the block, check its IP first. A fake ignores robots.txt by design, and our guide to stopping AI scraping covers blocking it at the firewall.

Does blocking CCBot remove pages already archived?

No. A robots.txt rule only affects crawls that start after CCBot reads it, and every snapshot already published keeps the pages it captured.

Common Crawl explained why in a November 2025 statement responding to an article in The Atlantic. Its archives are WARC files in an immutable format, which it says it cannot edit after publication without breaking their integrity. When a publisher asks for removal, it filters those URLs from later crawls and makes them inaccessible through its public index tools.

Copies are a separate problem. Anyone who downloaded a snapshot still has it, and datasets built from it keep their own copies. FineWeb's changelog shows its maintainers removing specific domains in January 2025 after a cease-and-desist notice, so each dataset handles removals on its own.

To ask Common Crawl about past content, email info@commoncrawl.org, the address its opt-out post gives. It records every legal opt-out request it receives in a public Opt-Out Ledger.

Common Crawl index: how to check if your site is in it

Search the index server at index.commoncrawl.org, which lists every capture of a URL pattern in a given crawl.

  1. Open the index server and pick a crawl, such as CC-MAIN-2026-39, the September 2026 crawl.
  2. Enter example.com/* for every captured URL on that host, or *.example.com to include subdomains.
  3. Read the results. Each capture shows a timestamp, the URL, the HTTP status, the MIME type and the WARC file that holds the copy.

The same query works from the command line and returns one JSON object per line.

curl "https://index.commoncrawl.org/CC-MAIN-2026-39-index?url=example.com/*&output=json&limit=20"

Check several crawls, since each one is a sample and the collection list holds 128 indexes back to 2008. Go slowly. The FAQ says the server is heavily rate-limited, asks you to sleep between calls and says a 503 means slow down. When we tested it in October 2026 it returned 504 timeouts for minutes at a time, so plan to retry. Treat an empty result as weak evidence, too, because Common Crawl says a "no captures" result reflects how its indexes are built, not what it stores.

Check whether your robots.txt blocks CCBot

Our robots.txt checker fetches your live robots.txt and checks whether 14 AI crawler tokens, CCBot included, may fetch the URL you enter. For each crawler it shows allowed or blocked, the rule that decided it, and whether the crawler fell back to your User-agent: * group. A blocked training crawler such as CCBot comes back as a note. A blocked AI search or user-triggered crawler comes back as a warning, because only those cost you citations.

The same run audits the file for syntax errors, conflicting rules and sitemap lines, and tests any path for any user agent you type, so you can enter CCBot and /premium/ and see which line matches. It reads robots.txt only and cannot prove that CCBot obeys it. A run costs 10 credits.

Keep reading