How to Get Cited by ChatGPT in 2026 (1.4 Million Prompts Analyzed)
Ahrefs analyzed 1.4 million ChatGPT prompts and found something most people building for AI search don't realize: ChatGPT retrieves dozens of URLs to answer a single prompt, then cites only about half of them. Roughly 23.4 million URLs got credit; roughly 23.4 million got read and ignored. The interesting question — the one I walk through in this video — is what separates the pages that get the citation from the pages that do the work anonymously.
I've been running my own experiments on AI citations with my clients for a while, so this study was a chance to check my playbook against a genuinely large dataset. Most of it held up. Some of it surprised me. Here's the full breakdown, with the numbers.
ChatGPT only credits half of what it reads
The headline stat first. Across the dataset — 1.4 million ChatGPT prompts from February 2025, analyzed with Ahrefs data scientist Xibeijia Guan — 50.02% of retrieved URLs were never cited. On an average prompt, ChatGPT pulls in roughly 16.57 URLs it will cite and 16.68 it won't.
So being retrieved is not the goal. Your page can feed the answer — ChatGPT can read it, learn from it, and paraphrase it — and your brand still gets zero visible credit. Getting cited is a second, separate contest, and it has its own rules.
The gatekeeping layer: your title, snippet, and URL get judged first
This is the part I find most actionable. According to research by AI search expert Dan Petrovic, referenced in the study, when ChatGPT retrieves search results each one comes back as a small bundle of metadata: the page title, a brief snippet or summary, the URL, and an ID number. ChatGPT uses that — not your actual content — to decide which pages are worth opening at all.
In other words, there's a gatekeeping layer before the model ever reads your page. Your title, your snippet, and your URL structure are doing the heavy lifting in that initial decision. If those three don't look relevant to the question being asked, your 3,000-word masterpiece never gets opened, let alone cited. I'd bet the same mechanic applies in some form to Claude, Perplexity and the rest — which is why title and URL discipline shows up in every AEO checklist worth reading.
The five retrieval channels — and why search dominates
Not every URL enters ChatGPT's system the same way. The study found an internal field called ref_type that labels the retrieval channel each URL came through, with five categories: search, news, Reddit, YouTube, and academia. The citation rates between them are wildly uneven.
The general search index dominates both in volume and citation rate — per the study, 88% of the URLs that end up cited by ChatGPT are taken directly from search. If you want to be cited, you need to be in that search selection pool, which means your content needs to be indexed and rank somewhere.
Here's the nuance from my own tests, though: it is not a one-to-one relationship with rankings. I've watched pages sitting on the tenth page of Google — positions no human ever clicks — get pulled in and cited because the content satisfied the search intent of the prompt. Ranking gets you into the pool. Intent match wins the citation. That matches what we see across the wider GEO statistics too: classic SEO and LLM visibility are correlated, but not identical.
Reddit: the textbook ChatGPT reads but never credits
The most striking finding in the dataset. Reddit has its own dedicated ref_type in ChatGPT's retrieval system — over 16 million data points in this study, apparently pulled in via a dedicated API feed on top of regular web search. Yet it gets cited at a rate of just 1.93%, and 67.8% of all non-cited URLs come from Reddit.
ChatGPT is using Reddit aggressively to understand topics, gauge consensus, and build context — then it almost never says so. For brands, the implication cuts both ways: Reddit shapes what the model believes about you even when no citation ever shows it, which is exactly why we watch Reddit threads inside our LLM visibility tracker — the influence is real even when the credit isn't.
Snippets and dates look like signals — they mostly aren't
This section is a masterclass in not fooling yourself with data. At first glance the aggregate numbers said something weird: non-cited pages had more metadata than cited ones — snippets 14.81% of the time versus 4.36%, and publication dates on roughly 93% of non-cited URLs versus 36% of cited ones. Less-optimized pages getting cited more? That would be a paradox.
It turned out to be a compositional artifact. The non-cited pool is overwhelmingly Reddit, and Reddit content pulled via API naturally carries date metadata. And per David McSweeney's research into ChatGPT's pipeline, the model actually abandons the snippet field once it decides to cite a URL — it opens the full page instead. Low snippet rates on cited pages are a byproduct of the plumbing, not a preference. Isolate the search channel only and the numbers collapse:
The honest takeaway is that no strong conclusion about snippets or dates survives once you account for where the URLs came from. Any citation research that compares cited versus non-cited without controlling for retrieval channel risks mistaking dataset quirks for ranking factors.
Titles that match the fan-out queries win the citation
Here's where the study confirms the thing I keep repeating to clients. When you prompt ChatGPT, it generates fan-out queries — internal sub-questions spun off your original prompt to hunt for specific facts. The study measured cosine similarity between page titles and those queries, and the pattern is clean: cited pages had a title similarity of 0.602 against the original prompt versus 0.484 for non-cited, and 0.656 against the best-matching fan-out query.
Within the search ref_type specifically, the gap gets even sharper. The practical rule: write titles that answer the question your buyer actually types into ChatGPT — and the sub-questions that prompt fans out into — rather than titles optimized for cleverness. This is the citation-side version of what prompt tracking gives you on the measurement side: the exact phrasing the models care about.
Two smaller levers: readable URLs and freshness
Two more findings worth acting on. First, URLs: search results with natural-language slugs had an 89.78% citation rate versus 81.11% for opaque ones. A slug like /best-ai-seo-tool carries semantic signal; /p?id=83629 carries nothing. SEOs have leaned on this rule for a decade; now there's citation data behind it.
Second, age. The median cited page in the search index is around 500 days old, with some cited pages over 2,700 days old — freshness is not a hard requirement for evergreen content. But it is a tiebreaker, and in the news channel it's decisive: cited news pages skew far younger.
If you publish commentary on moving topics, recency buys citations. If you publish evergreen comparisons and how-tos, relevance beats recency.
What I'd do with this data, in order
- Get indexed, then get rankable. 88% of citations come from the search pool. If Google can't find you, ChatGPT can't cite you. You don't need page one — you need to exist in the index with content that satisfies the intent.
- Rewrite titles against real prompts. Take the prompts your buyers type, list the fan-out questions inside them, and make sure a page title in your library matches each one semantically.
- Fix opaque URLs. Natural-language slugs are an 8-point citation-rate edge that costs you a redirect.
- Treat Reddit as influence, not citations. It shapes answers invisibly. Monitor the threads in your category; don't expect credit.
- Measure citations, not vibes. Track which prompts mention you, which pages get cited, and how that moves over time. That's the feedback loop everything above plugs into — it's exactly what we built our AI visibility tracker to do, and our AI SEO agent to act on.
FAQ
How do I get my website cited by ChatGPT?
Get into ChatGPT's search selection pool by being indexed and ranking for relevant queries, then win the selection step with a page title that semantically matches the user's prompt and its fan-out queries, plus a natural-language URL. In the Ahrefs study of 1.4M prompts, 88% of cited URLs came directly from the search channel, and cited pages had measurably more relevant titles (0.602 vs 0.484 cosine similarity).
Does ranking on page one of Google guarantee ChatGPT citations?
No. Ranking correlates with citations because search feeds the retrieval pool, but ChatGPT cites pages that satisfy the prompt's intent even when they rank far outside the top ten — I've seen pages on the tenth page of Google get cited in my own client work. Conversely, ChatGPT ignores roughly half the URLs it retrieves, so a ranking alone earns nothing.
Why does ChatGPT rarely cite Reddit?
Reddit enters ChatGPT's system through its own dedicated retrieval channel — over 16 million data points in the Ahrefs study — but only 1.93% of those URLs get cited, and 67.8% of all non-cited URLs are Reddit pages. The model uses Reddit to understand topics and gauge consensus, then attributes its answer to other sources.
Does content freshness matter for AI citations?
It depends on the channel. The median cited page in ChatGPT's search index is around 500 days old, and some cited pages are over seven years old — so evergreen content keeps earning citations. In the news channel, cited pages skew much younger (median around 200 days), so freshness matters most for time-sensitive topics.
How can I track whether ChatGPT cites my brand?
Run your buyers' actual prompts across the major models on a schedule and log which answers mention you, what sentiment they carry, and which URLs they cite. A monitor like Arvow's does this across ChatGPT, Claude, Gemini, Perplexity and Grok — see our guide on how to track ChatGPT brand mentions for the full setup.
Want to know which of your pages ChatGPT already cites — and which prompts you're invisible on? Start free with the LLM brand mention tracker, or run the full audit in the LLM visibility tracker.
The full study — Why ChatGPT Cites One Page Over Another — is worth reading end to end. But the one-line summary is almost old-fashioned: the pages that get cited are the ones whose titles and content match the questions the AI is asking behind the scenes, surfaced through the right retrieval channel. Write for the question, be findable, and the citations follow.
Generate, publish, syndicate and update articles automatically
The AI SEO Writer that Auto-Publishes to your Blog
- Cancel anytime
- Articles in 30 secs
- Plagiarism Free