GEOAI SEOChatGPT

How to Get Cited by ChatGPT: Crawler, Bing and Page Checklist

ChatGPT cites pages OAI-SearchBot can fetch, an index can find and a reader can quote. The crawler checklist, the Bing step, our 403 story and page traits.

Andrey Dereviankin

COO · AI Search & Data Architecture

Published Updated 15 min read
A dotted crawler path passes through an open gate, an index and a page, then ends at an AI answer card where one citation is highlighted in orange.

To get cited by ChatGPT, a page has to clear three gates in order. OpenAI’s search crawler, OAI-SearchBot, must be able to fetch it through your robots.txt and your CDN. The page must be findable in the indexes ChatGPT search reads: OpenAI’s own and its search partners’, where Bing is the one you can inspect and submit to. Then the page must state the answer more plainly, with more specific and better-sourced facts, than the other pages ChatGPT reads for that question. Mentions on the third-party pages ChatGPT already cites do the rest.

This guide is built on OpenAI’s, Microsoft’s and Cloudflare’s own documentation, checked on 3 October 2026, and on our tracker, which asks ChatGPT, Perplexity and Google the same buying questions every day.

What decides whether ChatGPT cites a page?

OpenAI’s position is short: “Any public website can appear in ChatGPT search,” and “Placement is not guaranteed” (OpenAI Help Center). There is no submission form and no published ranking formula. ChatGPT ranks results “using multiple factors intended to help users find relevant, reliable information.”

What you control is the order of the gates. A page the crawler cannot fetch never reaches the index. A page outside the index is never a candidate. A candidate that states nothing specific loses to one that does.

Our tracker shows how high the last gate sits. Asked “How do you get cited by ChatGPT?” every day from 4 September to 3 October 2026, ChatGPT gave 28 answers, and every link in them pointed to an OpenAI page: the crawler documentation, the publishers FAQ, the search help article, the launch post. On the same question Google’s AI Overviews and Perplexity cited Quora, Reddit, Search Engine Land and a dozen agency blogs. ChatGPT cited none of them.

For "How do you get cited by ChatGPT?" from 4 September to 3 October 2026, ChatGPT's 28 answers linked only OpenAI pages, while the Quora, Reddit and Search Engine Land pages that Perplexity and AI Overviews cited in 56, 53 and 46 answers got no ChatGPT link.

Same question, two source lists: ChatGPT linked only OpenAI’s pages, and the pages the other engines cited most got no ChatGPT link at all.

When a primary source can answer, ChatGPT goes to the primary source. A page that repeats the documentation gives it no reason to cite you. A page that adds what the documentation lacks, such as your own test, your numbers or your price, does.

Which OpenAI crawler has to reach your site?

OpenAI lists four user agents on its crawler page, and only one of them decides whether you can appear in ChatGPT search.

User agentWhat OpenAI uses it forWhat blocking it does
OAI-SearchBotSurfaces websites in ChatGPT’s search resultsYour pages “will not be shown in ChatGPT search answers”, though they can still appear as plain navigational links
GPTBotCrawls content that may be used to train OpenAI’s modelsOpts you out of training; no effect on search
ChatGPT-UserVisits a page when a user’s question or a custom GPT calls for itLittle: OpenAI says robots.txt rules “may not apply” to these user-initiated visits, and they do not decide search eligibility
OAI-AdsBotChecks pages submitted as ChatGPT adsMatters only if you advertise in ChatGPT

The most common mistake is allowing the wrong one. On 3 October 2026 Google’s AI Overview for “How do you get cited by ChatGPT?” told searchers to check robots.txt “to make sure GPTBot and Bing indexer can access your website”. GPTBot is the training crawler. Allowing it does nothing for citations, and blocking it costs nothing in ChatGPT search.

A site that wants citations but not training can say so in five lines:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

OpenAI says each setting “is independent of the others” and that a robots.txt change can take about 24 hours to reach search.

What is on the OAI-SearchBot checklist?

OpenAI’s help center sets two conditions for eligibility: allow OAI-SearchBot to crawl the site, and “confirm that the website host or content delivery network allows traffic from OpenAI’s published searchbot IP addresses.” In practice there are six places where the crawler can be stopped. Check each on the live site, not in your repository.

Diagram of OAI-SearchBot's route to a page through six checkpoints in request order: a link from a hub, the live robots.txt, CDN bot settings, the IP allowlist, a 200 response and the answer in the HTML, with the CDN checkpoint highlighted.

Six places OAI-SearchBot can be stopped on its way to a page. The CDN can block it even when robots.txt lets it in.

  1. The robots.txt visitors actually receive. Open yourdomain.com/robots.txt in a browser. A CDN can serve its own version: Cloudflare’s managed robots.txt adds Disallow rules for known AI crawlers, and ours carried a block our repository never contained.
  2. The CDN’s bot settings. “Block AI bots” switches, per-crawler blocks, bot-fight modes, WAF rules and rate limits can all return 403 to a crawler that robots.txt allows. Since July 2025 Cloudflare has asked every new domain whether to allow AI crawlers and blocks training crawlers unless the owner opts in (Cloudflare). Watch the training switch in particular: Cloudflare now says blocking AI training also blocks mixed-purpose crawlers, and it names Googlebot and Bingbot (Cloudflare). When we tried “Training: Block” on our own zone, the dashboard flagged Googlebot and Bingbot as blocked, and we reverted it within minutes.
  3. IP allowlists. If your firewall admits only known IPs, add the ranges OpenAI publishes for the search crawler at openai.com/searchbot.json.
  4. The response. Key pages should return 200 without a login, a cookie wall, a country block or a JavaScript challenge.
  5. The text in the HTML. OpenAI does not document whether OAI-SearchBot runs JavaScript, so put the answer in the HTML your server sends, not in a tab, an image or a script that fills the page after load.
  6. A path to the page. OpenAI’s own list of reasons a page is missing from its index includes “The site uses CDN or bot-blocking behavior”, “The page is new, rarely accessed, or low-signal” and “The site structure makes deeper pages harder to discover” (OpenAI Help Center). Link every page you want cited from your navigation or a hub page.

Then verify. The proof that OAI-SearchBot gets in is a 200 in your CDN or server log for a request whose user agent contains OAI-SearchBot (the current string ends in compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot) and whose IP is in OpenAI’s published list. A test from your laptop with a borrowed user agent proves something else, as we found out.

What happened when our own CDN blocked AI crawlers?

On the morning of 3 September 2026 we ran an access check on verticality.co with each AI crawler’s user agent. Every one of them got HTTP 403 from Cloudflare, including OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-User and GPTBot, while Googlebot, bingbot and an ordinary browser got 200 on the same pages in the same minute. Our robots.txt also carried rules we had never written: Cloudflare’s managed block, disallowing GPTBot, ClaudeBot, CCBot, Google-Extended and other training crawlers.

Cloudflare’s own log told a different story about the crawler that matters for ChatGPT. Over the previous 24 hours, OAI-SearchBot had 31 requests allowed and 5 blocked, and 3 of those 5 were our own test. The curl check had measured how Cloudflare treats a request that borrows a bot’s name, not how it treats the bot. The fetchers that really were locked out were Claude-User, with about 48 blocked requests, and Perplexity-User, with about 6 and none allowed. The cause was one setting: the legacy “Block AI bots” switch, set to block on all pages.

Chart from verticality.co, 3 September 2026: a test with borrowed user agents got 403 for every AI crawler, while Cloudflare's 24-hour log of real crawlers showed OAI-SearchBot allowed 31 times and blocked twice, Claude-User blocked about 48 times and Perplexity-User about 6 times with none allowed.

Our curl test said every AI crawler was blocked. Cloudflare’s log showed OAI-SearchBot getting through and Claude-User and Perplexity-User locked out.

We turned that switch off, cleared the per-crawler blocks and disabled the managed robots.txt the same afternoon, and a re-run returned 200 for every AI user agent. Two rules came out of it:

  • Test the real crawler, not an impostor. A request with a borrowed user agent shows how your CDN treats fakes. Real crawlers arrive from published IP ranges, so read the CDN’s per-crawler log before concluding anything.
  • Read the robots.txt you serve, not the one you wrote. Ours had gained a block of rules from a dashboard toggle.

Does ChatGPT search use Bing?

Partly, and less officially than most guides say. OpenAI’s help page says ChatGPT search “sometimes partners with other search providers” and rewrites your question into queries it sends them; for how those providers handle queries, it links two privacy statements, Microsoft’s and Shopify’s (OpenAI Help Center). The only explicit statement that Bing is ChatGPT’s default search is Microsoft’s, from May 2023 (Bing blog). OpenAI also runs its own crawler and refers to “OpenAI’s indexed and cached web content”.

So Bing is one input, not the whole pipe. It is also the only index outside OpenAI that you can inspect, submit to and get a report from, and Microsoft’s guidelines now say outright that Bing evaluates pages for “Copilot, and grounding API results”, not only for search (Bing Webmaster Guidelines).

We learned what that means in September 2026. Bing had every page of verticality.co in its index except two, /aeo-agency/ and /llm-seo-agency/, and Bing Webmaster Tools listed both as “Discovered but not crawled” on 24 September. ChatGPT had cited neither. Our GEO agency page, published the same week and indexed, was first cited by ChatGPT on 20 September, five days after launch, and was linked in 50 ChatGPT answers in the 30 days to 3 October. That fits Bing being an input; it does not prove it.

The two missing pages had something in common: they shared whole blocks of text. 17.8% of the five-word sequences on one also appeared on the other. Bing does not say why it leaves a page uncrawled, but its guidelines say “Duplicate URLs dilute signals and reduce Bing’s confidence in selecting a URL for grounding results or citations.” We gave each page its own angle, cut the shared text to 2.5%, requested indexing for both in Bing Webmaster Tools, pinged IndexNow and submitted the sitemap, which Bing had never processed. Being findable is the gate, not the finish line.

The Bing steps worth an hour:

  1. Add the site to Bing Webmaster Tools and inspect your key URLs. “Discovered but not crawled” means Bing knows the URL and has not fetched it. “Crawled but not indexed” means it fetched the page and chose not to keep it. They need different fixes.
  2. Submit the sitemap and check that Bing processed it. Ours sat unprocessed until we submitted it by hand.
  3. Send changed URLs through IndexNow. One request notifies Bing and the other participating engines (Naver, Seznam, Yandex, Yep), and a 200 response only means the URL was received (IndexNow). OpenAI is not a participant, so IndexNow speeds up Bing, not OAI-SearchBot.
  4. Keep pages distinct. Two pages that say the same thing in the same words compete for one slot.
  5. Check the AI Performance report. It counts citations in “Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations” (Bing blog). ChatGPT is not named, so treat it as a neighbouring signal, not a ChatGPT report.

What makes ChatGPT quote one page over another?

What ChatGPT lifts from our own pages shows what it looks for. When it describes Verticality, it repeats the fact rows at the top of our GEO agency page nearly word for word: who the service is for, what it covers, “$2,500–$7,500/month”. Each row is one complete, specific sentence ChatGPT can lift without editing. Five traits make a page work that way:

  • The answer comes first. The sentence under each heading answers the heading. Context follows.
  • Numbers carry a source and a date. “50 answers in 30 days, tracker data to 3 October 2026” can be cited; “lots of answers” cannot.
  • Every sentence names its subject. “Verticality’s GEO programme costs…”, not “it costs…”. A lifted sentence keeps no context.
  • Steps are lists and comparisons are tables. ChatGPT’s answers to how-to questions in our tracker are numbered lists, and a list on the page maps onto them.
  • The page shows when it was published and updated. ChatGPT’s choice of sources is re-run every time: in our tracker only 26.4% of the URLs ChatGPT cited for a question were cited again for the same question the next day (4–23 September 2026). A citation is a selection you keep winning, not a position you hold, so keep the dates and the facts current.
Two fact rows from Verticality's GEO agency page above ChatGPT's answer of 30 September 2026 that repeats them: B2B SaaS, DTC and regulated brands; Reddit threads, third-party citations, entity work and tracking.

ChatGPT repeats our GEO page’s fact rows almost word for word: who it is for and what it includes.

Controlled research points the same way. In the study that named generative engine optimization, from Princeton and partners, the best-performing edits were citing sources, adding quotations and adding statistics, with a 30–40% relative gain on the study’s main visibility measure, while keyword stuffing did not help (Aggarwal et al., KDD 2024). The test used a lab generative engine, not ChatGPT, so read it as direction rather than a forecast. For how to write each paragraph so it stands alone, see how SaaS pages get cited by AI.

Which off-site pages get a brand into ChatGPT answers?

ChatGPT leans on brand websites more than other AI engines do, and on community sites less. Across 56 buying questions about AI search that we tracked from 4 to 23 September 2026, ChatGPT cited reddit.com in 7.9% of its answers, linkedin.com in 1.2% and quora.com in none. Perplexity cited LinkedIn in 51.0% of its answers to the same questions; Google’s AI Overviews cited Reddit in 35.0%.

Bar chart of 56 AI-search buying questions, 4–23 September 2026: ChatGPT cited reddit.com in 7.9% of answers, linkedin.com in 1.2% and quora.com in 0%, against 30.9%, 51.0% and 3.6% for Perplexity and 35.0%, 41.7% and 11.8% for Google AI Overviews.

ChatGPT cites Reddit, LinkedIn and Quora far less often than Perplexity and Google’s AI Overviews do.

The Quora case is the sharpest. For “How do you get cited by ChatGPT?”, a Quora for Business guide was the most cited page on the question: 56 of 86 answers across Perplexity and Google’s AI Overviews. ChatGPT cited it zero times.

So off-site work for ChatGPT runs in a different order than for Google or Perplexity:

  1. Your own pages first. For buying questions ChatGPT cites vendor and agency sites far more often than the other engines do (our per-engine source split), so the fact rows on your own pages carry the most weight.
  2. The roundups and comparison articles ChatGPT already cites for your questions. Run your buyers’ questions, list the third-party pages in the citations, and work on being accurately included in those.
  3. Reddit, LinkedIn and Quora for the other engines. They earn citations in Perplexity and AI Overviews; in ChatGPT they rarely appear as links.

How do you check whether ChatGPT cites you?

One screenshot of one good answer is not a measurement. ChatGPT’s sources change from day to day, so checking needs a fixed method:

  • A fixed set of questions. 20 to 50 questions your buyers ask, in their words, asked the same way every time.
  • A schedule. Daily if you can, weekly at minimum, so one lucky day does not read as a trend.
  • A log of every answer. Whether your brand is named, which of your URLs are linked, and which other pages are cited instead.
  • Referral traffic. ChatGPT “automatically includes the UTM parameter utm_source=chatgpt.com in referral URLs” (OpenAI publishers FAQ), so cited links that get clicked show up under that source in your analytics.
  • Your crawler log. OAI-SearchBot requests to a new page are the first sign it can be cited at all.

Our tracker at ai.verticality.co runs this daily for ChatGPT, Perplexity and Google’s AI Overviews, and every tracker number in this guide comes from it. The metrics to report from that log, and how to roll them into one score, are in how to measure AI visibility. For a first read before you build that log, our free AI visibility checker asks 40+ buyer questions in ChatGPT and Google AI Overviews and emails the result within 24 hours.

Get your ChatGPT gates checked

If you want this done on your own site, the free AI visibility audit in our GEO programme starts here: crawler access through your CDN, your Bing index status, and the pages ChatGPT cites for your buyers’ questions instead of yours. For the same gates on Perplexity, Google’s AI Overviews, Claude and Copilot, see LLM SEO, model by model.

Frequently asked questions

Does blocking GPTBot stop ChatGPT from citing my site?

No. GPTBot is OpenAI's training crawler; ChatGPT search uses a different crawler, OAI-SearchBot, and OpenAI documents the two settings as independent. A site can disallow GPTBot in robots.txt and still be cited by ChatGPT, as long as OAI-SearchBot is allowed and the CDN lets it through.

Can you pay OpenAI to be cited by ChatGPT?

No. OpenAI says placement in ChatGPT search is not guaranteed, and that ads do not influence ChatGPT's answers: ads run on separate systems and advertisers cannot shape, rank or alter the responses. A ChatGPT ad buys a labelled unit, not a citation inside the answer.

Does ChatGPT search use Bing or Google?

OpenAI does not name an engine in its current help pages. It says ChatGPT search sometimes partners with other search providers and links Microsoft's and Shopify's privacy statements for them; Microsoft announced Bing as ChatGPT's default search in May 2023. OpenAI also runs its own crawler, OAI-SearchBot, and its own index. None of OpenAI's search help pages names Google.

How long does it take for ChatGPT to cite a new page?

It can take days once the page is crawled: ChatGPT first cited Verticality's GEO agency page on 20 September 2026, five days after it went live. A full programme takes longer; Verticality plans for first signals in 4–8 weeks, compounding by month 3–6.

How can I see traffic from ChatGPT in Google Analytics?

ChatGPT adds utm_source=chatgpt.com to the links it sends from search answers, so those visits show up under that source in GA4 and most analytics tools. Clicks are only part of the picture: many answers name or cite a brand without a visit, so citations also need to be tracked by asking fixed questions on a schedule.

Want your brand named in AI answers?

Our free AI visibility audit asks ChatGPT, Perplexity and Google AI Overviews the questions your buyers ask, shows who gets named instead of you and which sources they cite, and ends with a ranked work plan.

Get a free AI visibility audit

Not ready for a call? Run the free AI visibility checker: 40+ buyer questions in ChatGPT and Google AI Overviews, report by email within 24 hours.