How to Get Cited by ChatGPT and Perplexity: A Practical Guide

ChatGPT and Perplexity

To get cited by ChatGPT, your pages need three things: they must be in Bing’s index, your retrieval crawlers must be unblocked, and each page must answer one specific question completely inside its first 100 words. Perplexity needs all of that plus two more: recent updates, and mentions on third-party sites it already trusts. These are two different jobs, and treating “AI visibility” as one project is why most GEO tactics produce nothing.

The steps below run in order of leverage. The first two are plumbing; get them wrong and nothing else can work.

Why do ChatGPT and Perplexity cite completely different sources?

ChatGPT and Perplexity cite completely different

Because they use different retrieval systems and different selection biases.

ChatGPT does not search on every prompt. It answers from training weights unless something triggers browse mode, and when browse fires it pulls candidates from Bing’s index. If you’re not in Bing, you cannot be retrieved, regardless of how you rank on Google.

Perplexity searches on every query. It has the strongest recency bias of the major engines and leans heavily on community sources — Everything-PR’s June 2026 synthesis of citation-tracking studies put Reddit at roughly 20–24% of all Perplexity citations, the highest single-domain concentration measured on any AI engine. The same research found only about 11% of domains cited by ChatGPT also appear in Perplexity citations.

That overlap number is the whole strategic point. A SaaS brand can be cited on 18 of 25 prompts in ChatGPT and 2 of 25 in Perplexity with identical content. If your dashboard shows one blended “AI visibility” score, you are averaging away the only signal that tells you what to fix.

How do I get cited by ChatGPT and Perplexity? An 8-step checklist

ChatGPT and Perplexity? An 8-step checklist

1. Unblock the retrieval crawlers — and check Cloudflare before 15 September 2026

Every major provider now runs two crawler families: training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended) and retrieval crawlers that fetch at query time (OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, Claude-SearchBot). Blocking a training crawler keeps you out of future model weights. Blocking a retrieval crawler removes you from live answers entirely. Most teams blocked all of them in 2023 with one CDN toggle and never revisited it.

Two things to do this week. Open yourdomain.com/robots.txt and confirm no Disallow: / sits under any retrieval user-agent. Then open Cloudflare → Security → Bots, because a WAF rule overrides robots.txt. Cloudflare’s July 2026 changelog set new defaults from 15 September 2026: Training and Agent bots blocked on pages that display ads, Search allowed. The trap is in the fine print — multi-purpose crawlers including Googlebot, Applebot and Bingbot are evaluated under every applicable rule, so blocking Training blocks them too. If you run ads and you’re on Cloudflare, opt out in Security settings before the date or check your defaults after it.

2. Confirm you’re actually in Bing’s index

Open Bing Webmaster Tools, submit your sitemap, and check indexed page count against your actual page count. Plenty of B2B sites are fully indexed in Google and half-indexed in Bing because nobody ever looked. Enable IndexNow while you’re there — it pushes new and updated URLs to Bing directly instead of waiting for a crawl.

3. Rewrite the first 100 words of your five highest-intent pages

Pick the five pages tied to revenue: your category page, your two strongest comparison pages, your pricing explainer, your flagship guide. Delete the scene-setting opening. Replace it with a direct, complete, standalone answer to the question the page targets.

The test is simple. Cut those 100 words out, paste them somewhere with no context, and check whether they still form a correct and complete answer. If a reader needs the next paragraph to understand them, an LLM does too, and it will pick a competitor’s page that doesn’t have that problem.

4. Chunk each page into self-contained sections under question headers

Retrieval systems don’t ingest your page. They chunk it and score the chunks. A section that starts with “This is where it gets interesting” scores badly because the chunk carries no meaning alone.

Write H2s as the literal question a buyer would type, then make each section answer it fully without depending on what came before. Slightly repetitive top to bottom. Far more citable.

5. Put verifiable evidence directly into the body text

The Princeton and Georgia Tech GEO study (Aggarwal et al., KDD 2024) tested this experimentally: adding quotations, statistics, and cited sources raised source visibility in generative engines by up to roughly 40%. Not opinions — attributable facts with names attached.

For a SaaS founder this means publishing your own numbers. Median implementation time across your last 50 accounts. Actual API latency at p95. Nobody else on the internet has that data, which makes your page the only place it can be retrieved from.

6. Build the comparison pages your buyers actually prompt for

Writesonic’s March 2026 test of 50 prompts across ChatGPT’s newest models found that prompts containing a year, a price constraint, or an “X vs Y” structure triggered web search 100% of the time. Those are exactly the questions your buyers ask before a demo call.

So publish the pages that match: “[Your product] vs [competitor]”, “best [category] tools for [segment] in 2026”, “[category] pricing compared”. Write them honestly, including where you lose. A comparison page that never concedes anything reads as marketing to a reranker and to a human.

7. Get cited on pages that already get cited — not just on your own site

This is the step most teams skip, and it’s the highest-leverage one. The University of Toronto’s GEO study (arXiv, September 2025) found AI search exhibits a systematic and overwhelming bias toward earned media — third-party authoritative sources — over brand-owned and social content.

Read that as an instruction. Your owned blog is the weakest surface you can optimise. Go and find the pages already being cited when someone prompts for your category: the roundup on an industry publication, the G2 or Capterra category page, the Reddit thread in your buyers’ subreddit, the analyst brief. Then earn a place in them — a contributed piece, a corrected listing, an honest answer in the thread, a data point a journalist can quote. One inclusion in a page that gets cited on every prompt in your category is worth more than a quarter of blog posts.

8. Set a refresh cadence, then measure per engine

Perplexity’s recency bias means citations decay. Put your top ten pages on a genuine 60-day review — update the numbers, add what changed, then change the date. A modified timestamp on unchanged content fools nobody twice.

For measurement, segment referral traffic by source rather than lumping it into “AI”. Then read your server logs for OAI-SearchBot and PerplexityBot hits. Which pages they fetch most is your best available signal of what each engine considers citable on your domain.

What goes wrong when B2B teams try this?

Three failures, over and over.

They optimise the surface they control and ignore the one that gets cited. Forty new blog posts, perfect structure, zero movement — because the engines were pulling from a comparison site the team never touched. Owned media is necessary. It is rarely sufficient.

Their content is invisible to the crawler. A lot of B2B SaaS sites render key content client-side. Retrieval crawlers are far less patient than Googlebot: they tolerate fewer redirect hops and often don’t execute your JavaScript. Fetch your own page with curl and no JS. If the answer isn’t in the raw HTML, it isn’t citable.

They measure in aggregate. One blended AI visibility number hides the fact that you’re winning ChatGPT and absent from Perplexity, which are opposite problems with opposite fixes.

Frequently asked questions

How long does it take to get cited by ChatGPT or Perplexity?

Perplexity retrieves in real time, so new or updated content can appear in citations within days of being indexed. ChatGPT’s browse layer depends on Bing indexing, which typically adds one to three weeks. Structural fixes to existing indexed pages show up fastest; earning citations on third-party sites takes months.

Do I need llms.txt to get cited?

No major AI engine has confirmed using llms.txt as a retrieval signal. It costs almost nothing to publish, so publish one if you like, but do not treat it as a substitute for being in Bing’s index, unblocking retrieval crawlers, or restructuring your pages.

Should I block GPTBot to protect my content?

You can block GPTBot and stay fully citable, because training and retrieval are separate crawlers. Blocking GPTBot keeps your content out of future training runs while OAI-SearchBot continues to make you eligible for citations in ChatGPT search. Blocking OAI-SearchBot is the one that removes you from answers.

Does ranking #1 on Google guarantee an AI citation?

No. Google rank correlates with citation but doesn’t determine it, partly because ChatGPT’s browse layer draws from Bing rather than Google. Pages that rank well and are structured for extraction get cited. Pages that rank well and open with 200 words of preamble often don’t.

GSR

Girdhari Singh Rajpurohit

Founder of G2S Technology and a digital marketing consultant with 10+ years of experience across SEO, content, and lead generation — working with businesses from local clinics to SaaS companies, remotely across India.

Read the full story →

Ready to be the answer, not just a result?

Book a free strategy call and find out exactly where your visibility gap is.

Book a Free Strategy Call