Log file analysis for SEO: see which AI bots visit your site
·7 min read
Log file analysis means reading your server's access log to see which bots requested which URLs and what status code each one got back. For AI crawlers it is the only full record you own, because GPTBot, ClaudeBot and PerplexityBot never run your analytics script and Search Console reports on Google's crawlers alone. Pull a month of logs, count requests per bot and per status code with awk, check each bot's IPs against the list its operator publishes, then look for refusals, wasted crawl and pages no search bot reached.
This post is about crawlers. People who click through to your site from ChatGPT or Perplexity show up in analytics, and our guide to tracking visits from AI assistants covers them.
Where to find your access log
On your own server the access log is a text file on disk. On managed hosting and CDNs you download or export it.
| Setup | Where the access log is |
|---|---|
| Nginx, distribution or nginx.org packages | /var/log/nginx/access.log |
| Apache on Debian or Ubuntu | /var/log/apache2/access.log |
| Apache on RHEL, CentOS or Fedora | /var/log/httpd/access_log |
| cPanel hosting | Metrics > Raw Access, downloaded as .gz files |
| Cloudflare | Logpush, HTTP requests dataset |
| Amazon CloudFront | Standard logs in W3C format |
Those paths are defaults. nginx -T | grep access_log prints the ones Nginx loaded, and on Apache you search the config for CustomLog. Rotated days sit beside the live file as access.log.1, access.log.2.gz and so on, and every command below reads them all with zcat -f, which handles plain and gzipped files.
Cloudflare opened Logpush to every plan on September 30, 2026, with 25 GB a month included. From the HTTP requests dataset, push ClientIP, ClientRequestUserAgent, ClientRequestURI, EdgeResponseStatus and EdgeStartTimestamp. Use ClientRequestURI, because ClientRequestPath drops the query string.
CloudFront standard logs name their columns in a #Fields line. The IP is c-ip, the user agent cs(User-Agent) arrives URL-encoded, and the query string sits in cs-uri-query, apart from the path. Neither CDN writes the combined format, so adjust the field numbers in the awk commands below.
Behind a CDN or load balancer, your origin log records the proxy's address instead of the bot's, and every IP check in this post fails. If the first column is full of Cloudflare's addresses, that is what happened. Nginx's realip module fixes it with real_ip_header CF-Connecting-IP and a set_real_ip_from line per Cloudflare range.
What a combined access log line contains
A combined log line holds the client IP, the time, the request, the status code, the bytes sent, the referrer and the user agent, in that order. This sample is a GPTBot request from an address inside OpenAI's published range:
132.196.86.14 - - [08/Oct/2026:06:25:24 +0000] "GET /pricing HTTP/1.1" 200 18342 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"| Field | Value in the sample | awk field |
|---|---|---|
| Client IP | 132.196.86.14 | $1, split on spaces |
| Time | 08/Oct/2026:06:25:24 +0000 | $4 and $5, split on spaces |
| Request | GET /pricing HTTP/1.1 | $2, split on quotes |
| Status | 200 | $9, split on spaces |
| Bytes sent | 18342 | $10, split on spaces |
| Referrer | - |
$4, split on quotes |
| User agent | Mozilla/5.0 ... GPTBot/1.4 ... | $6, split on quotes |
Nginx calls this format combined and uses it when an access_log line names no format, and Apache's combined format has the same layout. The commands below split on double quotes with awk -F'"', because user agents contain spaces.
How to count AI bot traffic per crawler and status code
Put the bot names in one pattern, then have awk count the first match in each user agent. This is the core of server log analysis for SEO.
BOTS='gptbot|oai-searchbot|chatgpt-user|claudebot|claude-searchbot|claude-user|perplexitybot|perplexity-user|googlebot|bingbot|applebot|ccbot|bytespider|meta-externalagent|amazonbot'
# Requests per bot
zcat -f /var/log/nginx/access.log* |
awk -F'"' -v re="$BOTS" '{ ua = tolower($6) } match(ua, re) { print substr(ua, RSTART, RLENGTH) }' |
sort | uniq -c | sort -rn
# Requests per bot and status code
zcat -f /var/log/nginx/access.log* |
awk -F'"' -v re="$BOTS" '{ ua = tolower($6) } match(ua, re) { split($3, f, " "); print substr(ua, RSTART, RLENGTH), f[1] }' |
sort | uniq -c | sort -k2,2 -k1,1rnI use match() instead of the usual grep -o because GPTBot's user agent contains "gptbot" twice, in the token and in the URL, so grep -o counts every GPTBot request twice. ClaudeBot's contact address does the same. match() stops at the first hit on each line.
Our list of AI crawler user agents has every token worth adding to the pattern. For ClaudeBot by day, path and IP, see the ClaudeBot guide.
How to tell a real bot from a fake one
Check the IP. Anyone can send GPTBot's user agent, but only OpenAI controls the addresses on its list. These operators publish lists in the same JSON shape, a prefixes array of ipv4Prefix and ipv6Prefix entries.
| Operator | Bots | IP list |
|---|---|---|
| OpenAI | GPTBot, OAI-SearchBot, ChatGPT-User | openai.com/gptbot.json, searchbot.json, chatgpt-user.json |
| Anthropic | ClaudeBot and the other Claude bots | claude.com/crawling/bots.json |
| Perplexity | PerplexityBot, Perplexity-User | www.perplexity.com/perplexitybot.json, perplexity-user.json |
| Googlebot and its other crawlers | common-crawlers.json and related files |
To check every GPTBot IP in your log at once:
- Pull the unique IPs that claimed to be GPTBot.
- Download OpenAI's list.
- Test each IP against every prefix in the list.
zcat -f /var/log/nginx/access.log* |
awk -F'"' 'tolower($6) ~ /gptbot/ { split($1, a, " "); print a[1] }' |
sort -u > gptbot-ips.txt
curl -s https://openai.com/gptbot.json -o gptbot.json
python3 -c '
import ipaddress, json, sys
nets = [ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix"))
for p in json.load(open(sys.argv[1]))["prefixes"]]
for line in sys.stdin:
ip = ipaddress.ip_address(line.strip())
print(ip, "listed" if any(ip in n for n in nets) else "NOT LISTED")
' gptbot.json < gptbot-ips.txtSwap the bot name and list URL to check ClaudeBot or PerplexityBot. Block unlisted IPs by address, since a rule on the user agent would hit the real bot too.
For Googlebot logs, Google documents a DNS check that needs no list. A reverse lookup on the IP must return a hostname ending in googlebot.com, google.com or googleusercontent.com, and a forward lookup on that hostname must return the same IP. This loop runs both for every IP that claimed to be Googlebot.
zcat -f /var/log/nginx/access.log* |
awk -F'"' 'tolower($6) ~ /googlebot/ { split($1, a, " "); print a[1] }' | sort -u |
while read -r ip; do
name=$(host "$ip" | awk '/pointer/ { print $NF; exit }')
case "$name" in
*.googlebot.com.|*.google.com.|*.googleusercontent.com.)
if host "$name" | grep -qwF -- "$ip"; then echo "$ip verified $name"; else echo "$ip FAKE forward lookup fails"; fi ;;
*) echo "$ip FAKE ${name:-no reverse DNS}" ;;
esac
doneWhat to look for in your crawler logs
Once you know which requests are real, look for refusals sent to bots you meant to allow, crawl spent on parameter URLs and pages the search bots never reach.
403 and 429 responses sent to bots you want
When a verified GPTBot or Googlebot gets 403 or 429 responses, look at your firewall, your CDN's bot settings and your rate limits, because robots.txt never sends a status code. This command lists the refused requests by bot, status and path.
zcat -f /var/log/nginx/access.log* |
awk -F'"' -v re="$BOTS" '{ ua = tolower($6); split($3, f, " ") } f[1] ~ /^(401|403|429)$/ && match(ua, re) { split($2, r, " "); print substr(ua, RSTART, RLENGTH), f[1], r[2] }' |
sort | uniq -c | sort -rn | head -20Google's status code documentation says 429 and 5xx responses make its crawlers slow down for a while, and asks site owners not to use 401 or 403 to limit crawl rate. Other 4xx codes leave crawl rate alone, but Google won't index a URL that returns one.
A robots.txt block looks different in the log. A compliant bot shut out of the whole site requests /robots.txt, gets a 200 and fetches nothing else. If a bot you want behaves that way, run our robots.txt checker for AI crawlers, which names the rule that blocks a path for Googlebot or any crawler you name, and checks 14 AI crawler tokens against the URL you enter.
Crawl spent on parameter URLs
Filters, sort orders and tracking tags can turn one page into thousands of URLs. Measure how much bot traffic carries a query string, then find the parameters behind it.
# Share of bot requests with a query string
zcat -f /var/log/nginx/access.log* |
awk -F'"' -v re="$BOTS" 'tolower($6) ~ re { split($2, r, " "); print (index(r[2], "?") ? "query string" : "clean path") }' |
sort | uniq -c
# Most crawled parameter names
zcat -f /var/log/nginx/access.log* |
awk -F'"' -v re="$BOTS" 'tolower($6) ~ re { split($2, r, " "); print r[2] }' |
grep '?' | sed 's/^[^?]*?//' | tr '&' '\n' | cut -d= -f1 |
sort | uniq -c | sort -rn | head -20When one parameter dominates, stop linking to its variants from your own pages first. If bots keep fetching them, disallow the pattern in robots.txt with a line like Disallow: /*?*sort=.
Pages the search bots never reach
Compare your sitemap's paths with the paths search bots requested. What's left is every page no search crawler fetched during the log window.
curl -s https://example.com/sitemap.xml | grep -oE '<loc>[^<]+' |
sed -E 's#<loc>https?://[^/]+##' | sort -u > sitemap-paths.txt
zcat -f /var/log/nginx/access.log* |
awk -F'"' 'tolower($6) ~ /googlebot|bingbot|oai-searchbot|claude-searchbot|perplexitybot/ { split($2, r, " "); sub(/\?.*/, "", r[2]); print r[2] }' |
sort -u > crawled-paths.txt
comm -23 sitemap-paths.txt crawled-paths.txtVerify the IPs before you trust this list, since fake Googlebots add paths the real one never fetched. If your sitemap is an index, run the first command on each child sitemap. Paths must match exactly, so a trailing slash in one file and not the other will show up as a false gap.
Pages on that list often have few internal links pointing at them. Our orphan page finder crawls up to 150 pages from your homepage and compares them with your sitemap, so you can see which ones nothing links to.
Do you need Googlebot logs if you have crawl stats?
For Google's own crawling, mostly no. Search Console's Crawl Stats report groups Googlebot's requests by response, file type, purpose and Googlebot type, though only for root-level properties. It can't show which IPs borrowed Googlebot's name, and it has nothing on GPTBot or ClaudeBot. For those, the log is your only source.
How often to repeat log file analysis
Start with 30 days as a baseline, repeat the check monthly, and run it again within a week of any change to robots.txt, firewall rules, your CDN or your URL structure. Those are the changes that lock out a bot you want, and no dashboard tells you when it happens.
Check how long your logs are kept first. ls -l /var/log/nginx/ shows the oldest rotated file. If rotation deletes them sooner than 30 days, copy the files somewhere each week or switch to a CDN export.
Analyze AI bot traffic in your access log
Our AI Bot Log Analyzer runs the counting and IP checks from this post on a paste of up to 750,000 characters. It reads combined logs from Apache and Nginx, JSON lines from Cloudflare Logpush, Vercel or Caddy, and W3C logs from IIS or CloudFront. It checks each crawler's IP against the lists from OpenAI, Anthropic, Perplexity, Google, Bing, Apple, Common Crawl and DuckDuckGo, and when the IPs belong to a proxy it says so instead of calling real crawlers fakes.
The report splits requests per crawler into training, AI search and user fetches, and counts verified requests apart from ones that borrowed a crawler's name. It lists 401, 403 and 429 refusals, paths that returned errors, pages assistants fetched for a user, and requests by day. It doesn't compare against your sitemap, so keep the comm command for that.
A run costs 8 credits and never connects to your server. If your log already records the real client IP, filter it through grep -iE "$BOTS" before you paste, and 750,000 characters cover more days. Behind a proxy, paste unfiltered lines, because the report uses ordinary traffic to spot proxy addresses.