Bytespider: what ByteDance's crawler does and how to stop it

·7 min read

Bytespider is the web crawler run by ByteDance, the company behind TikTok and the Chinese news app Toutiao. ByteDance documents Bytespider only as the crawler for Toutiao Search, but reporters tie it to gathering training data for ByteDance's AI models, and reports from 2019, 2024 and 2026 say it ignores robots.txt. To stop it, add a User-agent: Bytespider group with Disallow: / to robots.txt and back it with a server or CDN rule that returns 403 to any user agent containing "bytespider".

Unless you want readers from Toutiao Search in China, blocking it costs you nothing we can find. The server rule does the real work, because it holds whether or not the bot reads your robots.txt.

Who runs Bytespider and what is it for?

ByteDance runs Bytespider, and the only documentation we found is in Chinese, on the Toutiao Search Webmaster Platform. Its help text says the Toutiao Search crawler's user agent is Bytespider, with a capital B. Site owners can submit a sitemap so Bytespider indexes their pages, set a daily cap on how many pages it crawls and write to zhanzhang@bytedance.com.

We found no English crawler page, no statement on robots.txt and no machine-readable IP list like the ones OpenAI and Anthropic publish. For crawler IPs, the help points to a Toutiao article from November 2019 that we could not load.

ByteDance has never said Bytespider collects AI training data. Others make that link. Fortune reported in October 2024 that the scraping looked aimed at ByteDance's next large language models, citing people familiar with the company. The bot directory Known Agents files it under AI data scrapers and hedges with "allegedly". For a blocking decision, treat it as a training crawler. The only visitors ByteDance says it brings you come from Toutiao Search.

The Bytespider user agent string

Bytespider names itself in the user agent, so a case-insensitive match on "bytespider" catches it. CrawlConsole records this form, which presents as an Android phone:

Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)

Known Agents lists a desktop Chrome version:

Mozilla/5.0 (compatible; Bytespider; spider-feedback@bytedance.com) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.0.0 Safari/537.36

The strings added to the crawler-user-agents project in 2019 are shorter and end in "Bytespider;bytespider@bytedance.com" or a bare "; Bytespider". Write your rules against the token, never the full string, or the next variant walks past them.

The string proves nothing on its own. With no IP list from ByteDance, a scraper that copies it looks exactly like the real crawler in your log, and a user agent rule blocks both.

Does Bytespider respect robots.txt?

Three reports say it does not, and ByteDance has published no robots.txt policy that we could find.

The complaints started in China. In November 2019, China Entrepreneur magazine reported, in a story republished by Huxiu, that webmasters saw Bytespider ignore robots.txt entirely and slow their sites. One small-site owner counted 460,000 requests in a single morning. ByteDance told the magazine the reports were inaccurate and pointed owners to Toutiao Search's email feedback channel.

In October 2024, Fortune cited research by the bot-management firm Kasada, which found that Bytespider "does not respect robots.txt". ByteDance and TikTok did not answer Fortune's emails.

The latest numbers come from TollBit. Its State of the Bots report for the first half of 2026, as Search Engine Journal summarized it on August 14, found that Bytespider reached disallowed pages on close to half of the European sites whose robots.txt named it. ChatGPT-User and Youbot did the same. TollBit counts any request to a disallowed URL as a bypass, whatever the operator claims.

One caveat applies to all three. Nobody outside ByteDance can separate real Bytespider requests from scrapers using its name, so some violations may belong to impostors. For your site the answer is the same. A robots.txt line may not stop traffic that calls itself Bytespider, and a server rule will.

Should you block Bytespider?

Yes, on almost any site that doesn't publish for readers in China. Toutiao is a Chinese-language app, and its search is the only thing ByteDance's documentation offers in return for the crawling.

Plenty of sites already block it. Cloudflare reported in July 2024 that Bytespider made more requests than any other AI crawler on its network, was the AI bot its customers blocked most often, and had accessed 40.40% of the websites Cloudflare protects. Its 2025 crawler review showed Bytespider's share of AI-only crawler traffic falling from 42% in May 2024 to about 7% in May 2025, with its request volume down 85%. It still crawls. It no longer dominates.

Blocking works forward only, so pages it already collected stay collected. If your Chinese-language site gets visits from Toutiao Search, keep it allowed and use the platform's crawl cap instead.

How to block Bytespider in robots.txt

Add a group for the Bytespider token that disallows everything.

User-agent: Bytespider
Disallow: /

Put it at the root of every host you serve, because crawlers read a separate robots.txt for blog.example.com and www.example.com. RFC 9309, the robots.txt standard, tells crawlers to match user agent lines case-insensitively, so the capitalization doesn't matter. Our robots.txt checker shows whether your live file blocks Bytespider and which rule matches a given path.

Before you add a server rule, you can test whether the bot visiting you honors the file. Publish the Disallow and change nothing else for two days. RFC 9309 asks crawlers not to rely on a cached copy for more than 24 hours, so that is a fair wait. If Bytespider fetched /robots.txt and kept requesting other pages, or never asked for the file at all, the traffic using that name on your site ignores it.

Keep the robots.txt group after you add the server rule. It is the signal a compliant crawler reads, which is why the rules below leave /robots.txt reachable.

How to block Bytespider at the server or CDN

Return a 403 to any request whose user agent contains "bytespider", except requests for /robots.txt. Use 403 rather than Nginx's 444, which drops the connection without a reply. Log tools read a 403 as a refusal, while ours counts a 444 as an error.

Nginx

The map goes in the http block, outside any server block.

map $http_user_agent $is_bytespider {
    default 0;
    ~*bytespider 1;
}

server {
    # your existing listen, server_name and root lines

    set $block_bot $is_bytespider;
    if ($uri = /robots.txt) {
        set $block_bot 0;
    }
    if ($block_bot) {
        return 403;
    }
}

Run nginx -t before you reload.

Apache

This works in .htaccess or the virtual host, with mod_rewrite enabled.

RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} bytespider [NC]
RewriteCond %{REQUEST_URI} !^/robots\.txt$
RewriteRule ^ - [F]

Cloudflare

Cloudflare's AI Crawl Control lists Bytespider by name, with ByteDance as the operator. Open AI Crawl Control, find Bytespider on the Crawlers tab and set it to Block. Without a paid plan, AI Crawl Control recognizes crawlers by user agent string only (Cloudflare docs), the same signal the server rules use.

To keep the robots.txt exception, write a WAF custom rule instead, with the Block action and this expression:

(lower(http.user_agent) contains "bytespider" and http.request.uri.path ne "/robots.txt")

Cloudflare's policies that block whole classes of AI crawlers, and rate limits for scrapers that don't name themselves, are in our guide to stopping AI scraping.

How to confirm Bytespider is blocked

Search your access log for the token and check the status codes. After the server rule goes live, every Bytespider request should get a 403 except fetches of /robots.txt. If you've never pulled an access log, our log file analysis guide shows where hosts keep it. A refused request looks like this, with a placeholder IP:

203.0.113.24 - - [09/Oct/2026:06:41:12 +0000] "GET /pricing HTTP/1.1" 403 153 "-" "Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)"

These commands assume the Nginx or Apache combined format. The second one reads only the current file, so lines from before your change don't count.

# Status codes Bytespider received
zgrep -ih "bytespider" /var/log/nginx/access.log* | awk '{print $9}' | sort | uniq -c | sort -rn

# Pages it still got with a 200, robots.txt excluded
grep -i "bytespider" /var/log/nginx/access.log | awk '$9 == 200 && $7 != "/robots.txt" {print $7}' | sort | uniq -c | sort -rn | head -20

# robots.txt fetches per day
zgrep -ih "bytespider" /var/log/nginx/access.log* | awk '$7 == "/robots.txt" {print substr($4, 2, 11)}' | sort | uniq -c

Only 403s, plus robots.txt. The block works, unless a CDN serves cached HTML in front of it. Cached hits never reach your origin rule or your origin log, so block at the CDN in that case.

200s on other pages. The rule isn't matching. Check that it sits in the server block for the hostname Bytespider requests, and on Apache that AllowOverride lets .htaccess use mod_rewrite.

No Bytespider lines at all. That is what a Cloudflare block looks like from the origin, since the requests stop at the edge. Check AI Crawl Control instead. Its Crawlers tab shows allowed and unsuccessful requests per crawler, and its Robots.txt tab lists crawlers that requested paths your robots.txt disallows (Cloudflare changelog).

Check whether Bytespider still reaches your site

Our AI Bot Log Analyzer reads an access log you paste, up to 750,000 characters, in Apache or Nginx combined format, JSON lines from Cloudflare Logpush, Vercel or Caddy, or W3C format from IIS and CloudFront with its #Fields line. It recognizes Bytespider as ByteDance's training crawler, alongside GPTBot, ClaudeBot and the other AI bots it knows. ByteDance publishes no IP list, so the report counts Bytespider's requests and marks them "No list" in the Verified column instead of guessing whether they are real.

For each crawler you get its request count, the requests your server refused with a 401, 403 or 429, and the requests that hit other errors. The report also lists the paths crawlers requested most, the paths that returned errors and requests by day. Run it on a log from before the block and one from after, and in the second run every Bytespider request except the robots.txt fetches should count as refused. Nothing connects to your server, and a run costs 8 credits.

Keep reading