How to get cited by ChatGPT
·7 min read
ChatGPT citations come from ChatGPT search. A page gets cited when OAI-SearchBot is allowed to crawl it, its answer is in the HTML rather than loaded by JavaScript, and it comes back for the short search queries ChatGPT writes from the user's prompt. Of the pages that clear those bars, the ones quoted most tend to state a specific fact in a sentence that makes sense on its own.
GPTBot has nothing to do with it. That crawler collects training data, and blocking it does not remove you from ChatGPT's search answers. ChatGPT can name a brand it learned about in training without linking anything. A citation with your URL happens only when ChatGPT searches.
How does ChatGPT choose sources?
OpenAI does not publish a ranking formula, but its ChatGPT search help article describes the retrieval. ChatGPT rewrites the prompt into one or more targeted search queries, sometimes sends them to third-party search providers, reads what comes back and may send narrower follow-up queries before it writes the answer. OpenAI's own example turns a question about cancer drugs into "CCR8 immunotherapy drug development 2025", then "CHS-114 conference 2025".
So you compete twice. First the page has to come back for a query that reads like a terse keyword search. That part is ordinary SEO. Then the model has to find a passage on the page worth quoting, and that part comes down to how you write.
Google's AI Overviews choose sources from Google's own index, so they get their own post on how to show up in Google AI Overviews.
OAI-SearchBot vs GPTBot vs ChatGPT-User
OAI-SearchBot is the only OpenAI crawler that decides whether your pages can appear in ChatGPT search. OpenAI lists its crawlers on its bots page, and three of them matter here.
| User agent | What it does | Obeys robots.txt | Affects citations |
|---|---|---|---|
| OAI-SearchBot | Crawls pages so they can appear in ChatGPT search answers | Yes | Yes. Sites that block it are not shown in search answers |
| GPTBot | Collects content to train OpenAI's foundation models | Yes | No |
| ChatGPT-User | Fetches a page when a user's action in ChatGPT or a custom GPT asks for it | Not always, OpenAI says | No. OpenAI says it is not used to decide what appears in search |
You opt in or out with a robots.txt group for each user agent, and each setting is independent. OpenAI says search picks up a robots.txt change in about 24 hours. A site that blocks OAI-SearchBot can still show up as a plain navigational link, just never as a source inside an answer.
The full user agent string includes a version number, as in compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot. Match on the OAI-SearchBot token in firewall rules and log filters, and check that real requests come from the IP ranges OpenAI publishes at https://openai.com/searchbot.json.
Allow OAI-SearchBot in robots.txt
Give OAI-SearchBot its own group in robots.txt and allow the pages you want cited. This version lets ChatGPT search in and keeps training out:
User-agent: OAI-SearchBot
Allow: /
Disallow: /cart/
User-agent: GPTBot
Disallow: /
User-agent: *
Disallow: /cart/Notice the /cart/ line appears twice. A crawler that finds a group with its own name follows only that group and ignores User-agent: *, so every rule you still want for OAI-SearchBot has to be copied into its group. Leave it out and OAI-SearchBot may crawl paths you closed to everyone else. The robots.txt checker shows which rule matches OAI-SearchBot for any path.
Robots.txt is only half the job. If your CDN or firewall challenges unknown bots, OAI-SearchBot gets a challenge page instead of your content, so allow requests from the IP ranges above, as OpenAI asks publishers to do. A CDN setting that blocks AI crawlers as a group can override everything in your robots.txt without telling you.
Make the page readable without JavaScript
Put the answer in the HTML your server sends, because you cannot count on OpenAI's crawlers running your scripts. When Vercel and MERJ studied AI crawler traffic in December 2024, none of the major AI crawlers rendered JavaScript, and that included OAI-SearchBot. OpenAI's docs do not say whether that has changed, so build for a crawler that reads raw HTML only.
Test a page with curl, which fetches the raw HTML and runs no scripts:
curl -s -A "OAI-SearchBot" https://example.com/pricing | grep -i "per month"No output means the text is missing from the raw HTML or something blocked the request. Run it again without the grep to see which. Content hidden in tabs or accordions with CSS is fine, because it is in the HTML. Content fetched after a click is not. AI Crawler View compares the raw HTML with the rendered page and lists the headings, prices, structured data and links that appear only after JavaScript runs.
Write passages ChatGPT can quote
Open each section with one or two sentences that answer a single question with a specific fact, and name the subject in them instead of writing "it" or "our product". The model may lift that passage without the rest of the page, so it has to stand alone.
Compare two lines from a pricing page:
Weak: Flexible plans that grow with your team.
Strong: Acme Time costs $12 per user per month on the Team plan, billed yearly, with a 14-day trial.Nobody can cite the first line for anything. The second answers "how much does Acme Time cost" with the product name and the price in one sentence.
A page full of positioning copy can rank well and still never get cited, because no sentence in it is specific enough to quote. A few habits help:
- Write headings in the words people search with, then answer directly under them.
- Put comparisons, specs and plan limits in tables.
- Give numbers with their units and say where each one comes from.
Get your brand named in ChatGPT answers
To name you, ChatGPT has to know which company you are and find you on the pages it reads for your topic.
Make your brand easy to identify
Use one company name everywhere, in page titles, Organization schema, social profiles and directory listings. Say what you are in one plain sentence on the home page and the About page, such as "Acme Time is time-tracking software for construction crews." That is the sentence you want repeated back to people. Then mark up the site with Organization schema and list your official profiles in sameAs:
{
"@context": "https://schema.org",
"@type": "Organization",
"name": "Acme Time",
"url": "https://acmetime.example",
"logo": "https://acmetime.example/logo.png",
"sameAs": [
"https://www.linkedin.com/company/acmetime",
"https://github.com/acmetime"
]
}The brand entity checker tests whether one name appears across your Organization and WebSite markup, og:site_name, web manifest and llms.txt, whether your sameAs profiles exist, and whether a Wikidata item lists your site.
Earn mentions on the sites ChatGPT already cites
For buying questions like "best time tracking app for contractors", the sources are often comparison articles, review sites and forum threads rather than vendor pages. Find out which ones ChatGPT cites for your topic first, because the list differs a lot between categories.
Email the authors of cited comparison articles with accurate, current facts about your product, and ask for a correction if they list you with an old price. Claim your profile on any cited review site and ask real customers to review you there. Answer cited forum threads under your own name, say that you work for the company, and actually help. Moderators delete planted answers.
Keep cited pages current
Update the pages that already get cited first, and keep their numbers and dates true. A wrong price on a page ChatGPT quotes gets repeated to everyone who asks. Rewritten queries can carry a year, as in OpenAI's example, so a comparison page still titled "2024" looks stale next to one that says 2026.
Review cited pages whenever pricing, features or plans change, and at least once a quarter otherwise. Show a visible "Updated" date and set dateModified in the page's Article schema to the same day. Change the date only when the content changed.
How to track ChatGPT citations in analytics
You cannot see ChatGPT citations in any OpenAI report, but you can see the clicks they send. ChatGPT adds utm_source=chatgpt.com to the links in its answers, so those visits carry a source even when the app sends no referrer.
In GA4, open Reports, then Acquisition, then Traffic acquisition. Set the dimension to Session source and look for chatgpt.com. Older links can arrive as chat.openai.com, so match both domains, and expect a share of ChatGPT visits to land in Direct anyway. The AI referral tracking guide has filters ready to paste into GA4, Looker Studio, Plausible, PostHog and Matomo.
Clicks undercount citations, since many readers take the answer and never click. Server logs fill in part of the gap. OAI-SearchBot requests show which pages OpenAI is crawling for search, and ChatGPT-User requests show pages that a person's ChatGPT session opened.
Find the pages ChatGPT cites for your topic
AI Citation Finder shows which pages ChatGPT and Google AI Overviews cite. Enter a topic and it lists the 25 pages cited in the most AI answers, with the number of answers citing each one and the monthly searches behind them, plus the sites cited most and the share of answers each appears in. Enter a domain, yours or a competitor's, and it reports how many AI answers cite that site, which of its pages are among the 25 most cited, and which other pages those answers cite next to them. You can count ChatGPT, AI Overviews or both, in any of 20 countries or all of them together.
Our citation data tracks the questions people ask most, and far fewer ChatGPT answers than AI Overviews. Read the counts as a comparison between pages and sites, not as every answer ChatGPT has given. The tool does not read what the answers say and does not cover Perplexity or Gemini. A run costs 160 credits, and the cited sites it returns are your outreach list for the mentions work above.