GEO SEO checklist: sitewide and off-site checks with a test for each

·9 min read

A GEO SEO checklist covers the sitewide and off-site work that decides whether ChatGPT, Perplexity, Claude and Google's AI features can reach your site, read it and name your brand. Check that robots.txt lets the AI search crawlers in, that your CDN is not blocking them anyway, that your key content is in the raw HTML, that your brand reads as one entity, that the sites AI answers cite mention you, and that you track AI impressions and referrals. llms.txt is optional, because no major AI search engine has confirmed it uses the file.

Each item below has a test you can run in minutes. This list is sitewide. For per-page work such as question headings and answer-first sections, use the AEO checklist. For the background, read what generative engine optimization is.

The GEO SEO checklist at a glance

Item Passes when
AI search crawlers allowed No Disallow in your live robots.txt matches OAI-SearchBot, Claude-SearchBot or PerplexityBot on pages you want cited
CDN not overriding robots.txt Your CDN's AI bot settings allow search and agent bots, and your logs show no 403 or 429 responses to them
Content in the raw HTML curl output for a key page contains its price, headings and main answer
llms.txt, if you publish one The file returns 200 and every link in it loads
One brand entity The home page has Organization markup with a logo and sameAs, and Wikidata has an item that lists your site
Off-site presence Your brand appears on some of the domains AI answers cite for your buyers' questions
Measurement You have a monthly baseline of Google AI impressions and AI referral sessions

Which AI crawlers should your robots.txt allow?

Allow the AI search bots and the user-triggered fetchers on every page you want cited, and treat the training bots as a separate decision. Each vendor gives these jobs separate tokens, so you can block training and stay in search.

Token Operator What the operator says it does Usual setting
OAI-SearchBot OpenAI Surfaces sites in ChatGPT search. Opted-out sites don't appear in ChatGPT search answers Allow
ChatGPT-User OpenAI Fetches pages for user actions in ChatGPT and custom GPTs. Robots.txt rules may not apply Allow
GPTBot OpenAI Crawls content that may be used to train generative AI foundation models Your call
PerplexityBot Perplexity Surfaces and links sites in Perplexity search results. Not used to train foundation models Allow
Claude-SearchBot Anthropic Improves Claude's search results. Blocking it may reduce your visibility there Allow
ClaudeBot Anthropic Collects content that could contribute to model training Your call
Google-Extended Google Controls use of content for training future Gemini models and for grounding in Gemini Apps and Vertex AI Your call

The rows come from the crawler docs published by OpenAI, Perplexity, Anthropic and Google. OpenAI says its systems take about 24 hours to pick up a robots.txt change.

Google-Extended is the one people misread. Google says it does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode use pages Googlebot indexes, with no technical requirement beyond being indexed and eligible for a snippet, per Google's AI features guide. A site-level opt-out for those features is the Search generative AI setting in Search Console, which defaults to include.

This file blocks training and leaves search alone:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /cart/
Disallow: /search

The search bots have no group here, so they follow the * rules. That holds until someone adds Disallow: / under User-agent: * to stop scrapers and drops the site from ChatGPT search, Claude search and Perplexity at once. Under RFC 9309, a crawler uses the * group only when no group names it, so if you give the search bots a group of their own, copy in any paths you still want blocked. The ClaudeBot guide covers the Anthropic tokens and how to verify their IPs.

ChatGPT-User and Perplexity-User fetch pages because a person asked, and both vendors say robots.txt may not apply to them. Stopping them takes a firewall rule, and a firewall rule can also stop them by accident.

To test, run curl -s https://example.com/robots.txt, find the group each search token falls into, and confirm no Disallow covers the pages you want cited.

Is your CDN blocking AI crawlers that robots.txt allows?

Assume it might be until you check. A CDN or firewall rule can refuse the request before robots.txt matters, and the file shows no sign of it.

Cloudflare is the common case. Its AI bot policies sort bots into Search, Agent and Training. Agent means activity on a person's behalf, such as chat fetch bots, and Training includes crawlers that also do search. Each class can be allowed, blocked everywhere or blocked only on pages with ads. Since September 15, 2026, Cloudflare offers new domains a preset based on whether the site earns ad revenue. The ad-supported preset allows Search, sets Training to Disallow AI Training and blocks Agent bots on pages with ads. An ad-supported site added to Cloudflare since then may refuse live chat fetches that nobody chose to block.

With managed robots.txt on, Cloudflare also prepends its own block to your file. In Cloudflare's example, that block disallows GPTBot, ClaudeBot, Google-Extended and other training crawlers. Cloudflare calls robots.txt compliance voluntary and points to AI Crawl Control for per-crawler rules.

To test it:

  1. Compare your live robots.txt with the file you deploy. Groups at the top that you never wrote come from the CDN.
  2. Read the AI bot settings in your CDN dashboard. On Cloudflare they are in Security Settings.
  3. Count the status codes AI bots received in your access log.
grep -iE "oai-searchbot|chatgpt-user|perplexitybot|claude-searchbot|claude-user" access.log | awk '{print $9}' | sort | uniq -c

The awk field assumes the standard Nginx or Apache combined format. Any 403 or 429 there is a rule to find. Curling your site with a crawler's user agent proves little, because many bot rules check the operator's IP ranges, so your request can pass while the real bot is blocked.

Can AI crawlers read your content without JavaScript?

Most AI crawlers read only the HTML your server returns, so anything you want quoted must be in it. In a December 2024 study, Vercel and MERJ found that none of the major AI crawlers rendered JavaScript, GPTBot, ClaudeBot and PerplexityBot included. The exceptions were Gemini, which uses Googlebot's infrastructure, and Applebot.

curl -s https://example.com/pricing | grep -c "149"

Swap in a key page and a string from it, such as a price or the first sentence of your main answer. A count of 0 means the text arrives by JavaScript, and the fix is server-side rendering or static generation.

Does llms.txt help GEO SEO?

No major AI search engine has said it uses llms.txt, so it is optional. Google's AI optimization guide says its generative AI search needs no AI text files and that Search ignores llms.txt. The OpenAI, Anthropic and Perplexity crawler docs linked above do not mention it.

Publish one if coding agents read your docs and it takes you an hour, but only after the items above. Then curl -sI https://example.com/llms.txt should return 200 and every link inside should load. Our guide to what llms.txt is has a full example.

How do you make your brand one clear entity?

Use one brand name everywhere, put Organization markup with your logo and profile links on the home page, and get a Wikidata item that points to your site once independent sources have covered you.

Google's Organization structured data docs say the markup helps Google tell your organization apart from others. The logo must be at least 112 by 112 pixels and crawlable, sameAs lists your profiles on other sites, and the markup goes on the home page or an about page.

{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Example Co",
  "url": "https://example.com",
  "logo": "https://example.com/logo-512.png",
  "sameAs": [
    "https://www.linkedin.com/company/example-co",
    "https://github.com/example-co"
  ]
}

Google's AI optimization guide says structured data isn't required for its generative AI search. Add the block anyway. It states your name, logo and profiles in a form parsers read without guessing, and it takes ten minutes.

Wikidata's notability policy accepts items for entities that are clearly identifiable and described by serious, publicly available references. Set the official website property, P856, to your home page. Without independent coverage the item has nothing to cite, so do the next item first.

Names drift. "Example Co" in your markup and "ExampleHQ" on LinkedIn read as two candidates, so pick one and fix the other.

To test, view source on the home page, find one Organization block with name, url, logo and sameAs, then search Wikidata for your domain.

Which sites do AI answers cite in your niche?

Ask your buyers' questions with web search on and list the cited domains, because those sites shape what AI answers say about you.

  1. Write 10 questions a buyer asks before choosing a product like yours, such as "best payroll software for a 20-person company".
  2. Ask each one in ChatGPT and Perplexity with search on, and in Google's AI Mode.
  3. Record every cited domain and whether the answer names your brand.
  4. Sort the domains by how many answers cite them.

The top of that list is your outreach plan. Get listed or reviewed on those sites, and ask for corrections where they describe you wrongly. Answers vary between runs, so repeat the sample monthly. You pass when your brand appears in some answers and each top-cited domain that leaves you out has a plan.

How do you measure GEO results?

Track Google's AI impressions in Search Console and AI referral sessions in analytics, knowing neither covers everything.

Search Console's Generative AI performance report opened to all sites on August 31, 2026. It shows impressions only, by page, country, device and date, with AI Overviews and AI Mode combined. Google's AI features guide says traffic from those features also counts in the regular Performance report under Web search.

In analytics, split out sessions from chatgpt.com, perplexity.ai, gemini.google.com and claude.ai. ChatGPT adds utm_source=chatgpt.com to the links it cites, so check UTM sources too. Our AI referral filters have copy-paste setups for GA4, Plausible, Matomo and others.

You pass when both numbers sit in a monthly report with a baseline from before your changes.

Check your site against this GEO checklist

Our GEO tools cover this list item by item, priced per run. AI Crawler View compares a page's raw and rendered HTML and lists what appears only after JavaScript, for 8 credits. AI Bot Log Analyzer reads a pasted access log and shows which AI crawlers visited, which requests got a 401, 403 or 429, and visits that used a crawler's name from an IP its operator does not publish, for 8 credits. Brand Entity Checker tests one name across your markup, og:site_name, manifest and llms.txt, along with your logo, sameAs profiles and Wikidata item, for 8 credits. AI Agent Readiness validates llms.txt syntax and link URLs for 10 credits.

Off-site, AI Citation Finder lists the 25 pages Google's AI Overviews and ChatGPT cite most for a topic, and the sites cited most, for 160 credits. AI Answer Visibility asks ChatGPT, Perplexity and Gemini a buyer question with web search on and reports whether each answer names or cites you and who it names instead, for 120 credits. Treat both as samples of what users see. Google says no third-party tool has access to its ranking or AI systems, ours included.

For robots.txt, the Robots.txt and AI Crawler Checker in our SEO tools shows the rule each AI crawler matches for a path and the exact line to change, for 10 credits. It also reports when your server refuses its own request with a 401, 403 or 429.

Keep reading