AI Tools

Getting Cited by ChatGPT and Perplexity: How Writing Changes for LLM Visibility

Every Monday since February, the same 40 prompts have gone into ChatGPT and Perplexity from this desk. Buying questions, comparison questions, troubleshooting questions, all logged in a spreadsheet alongside every source each engine cited that week. The pattern that emerged was uncomfortable for anyone raised on classic SEO. Pages that had held Google's top three for years never appeared in the answers. Meanwhile a mid-tier vendor blog with a dated stat table got named week after week, and a two-year-old forum thread outperformed a 4,000-word pillar page on the same topic.

The spreadsheet kept pointing at one conclusion: these engines do not reward pages, they reward passages. Once that clicks, most of the old writing habits start looking like liabilities. This piece lays out what the research and six months of tracking agree on, and what it means for how content gets written now.

WHAT ACTUALLY CHANGES

• Passages beat pages. Engines retrieve whole pages but cite individual chunks, so every section has to stand on its own.

• Evidence earns citations. Princeton's GEO study measured up to 40% visibility gains from adding statistics, quotes and named sources.

• The two engines barely overlap. Only about 11% of domains cited by ChatGPT are also cited by Perplexity.

• SEO is the entry ticket, not the prize. Retrieval pools are built from search indexes, but selection happens on structure and specificity.

• Volatility is the norm. Reddit went from roughly 60% of ChatGPT responses to 10% in about six weeks. Nothing here is set-and-forget.

Four numbers that explain the whole game

The AI citation research space matured fast between 2024 and 2026. Several large, independent datasets now exist, and they converge on the same story: visibility inside AI answers follows its own rules, and those rules are measurable.

NumberWhat it measuresSource
40%Maximum visibility lift from GEO tactics like adding statistics and cited sourcesPrinceton, KDD 2024
11%Domain overlap between ChatGPT and Perplexity citations across 680M analyzed citationsAveri, 2026
21.9Average citations per Perplexity answer, versus roughly 8.3 for ChatGPTWhitehat SEO, 2026
2.3xMore citations for self-contained 50-150 word chunks than unstructured textZipTie.dev, 2026

How ChatGPT and Perplexity actually pick sources

Neither engine reads a prompt and searches for it verbatim. When ChatGPT decides a question needs live information, it breaks the prompt into multiple sub-queries, a process the industry calls query fan-out, and sends them to Bing and its own retrieval systems (Ahrefs, 2026). Perplexity runs a comparable pipeline against its own hybrid index and cites far more densely, attaching sources to individual claims rather than to the answer as a whole.

What happens next is the part most writers underestimate. The engines do not evaluate whole pages. Retrieved pages get split into chunks, chunks get converted into vectors, and the passages most semantically similar to each sub-query get pulled into the answer. Testing by SEO researcher Metehan Yeşilyurt puts ChatGPT's retrieval window at roughly 38 to 65 sources per search, and the vast majority of retrieved pages never earn a citation because the final selection binds to specific sentences (SERPsecrets, 2026).

1.  Query fan-out. One prompt becomes several sub-queries. A page ranking for "best CRM" can lose to a page matching the sub-query "CRM for two-person real estate teams".

2.  Retrieval. Search indexes assemble a candidate pool. Classic SEO signals still gate this stage: unindexed or blocked pages score zero.

3.  Chunking and selection. Pages get split into passages and compared by semantic similarity. Claim density, entity clarity and recency decide which chunks survive.

4.  Attribution. The winning passages get cited inline. One important quirk: ChatGPT deduplicates by domain, so it picks a single best page per site per claim.

The two engines apply this pipeline to very different source pools, which is why optimizing for one does not automatically cover the other.

SignalChatGPTPerplexity
Retrieval backboneBing plus proprietary systems, with query fan-out into sub-queriesOwn hybrid index with real-time retrieval
Citations per answer~8.3 on average~21.9 on average, the densest of the major engines
Favored sourcesWikipedia and encyclopedic references, LinkedIn (about 14.3% of responses per Semrush's 325K-prompt study), institutional sitesReddit and community discussion (roughly 20% concentration per Evertune), tier-one journalism, LinkedIn at 5.3%
Freshness sensitivityModerate; leans on established authorityHigh; one 2026 analysis found strong citation rates for content under 30 days old
Overlap with each otherOnly ~11-12% of cited domains appear on both enginesPer Averi's 680M-citation analysis and Passionfruit's cross-platform check

What the Princeton GEO study actually proved

The academic anchor for all of this is "GEO: Generative Engine Optimization" by Aggarwal and colleagues, presented at KDD 2024 and available as arXiv preprint 2311.09735. The team built a benchmark of 10,000 queries across multiple domains, tested nine content modifications on a system built to mimic Bing Chat, and validated the strongest tactics on Perplexity.

The results split cleanly into tactics that moved citations and tactics that did nothing or backfired.

TacticWhat it means in practiceMeasured effect
Statistics additionReplacing vague claims with specific, sourced numbersup to +40%
Quotation additionAdding attributed quotes from credible voices+30 to 40%
Cite sourcesLinking claims to named external references~+28% and up
Fluency optimizationRewriting for clearer, easier-to-parse prose, no new facts added~+28%
Authoritative voiceConfident, direct phrasing instead of hedgingpositive
Keyword stuffingAdding more target keywords, the classic SEO reflexno gain

Combined tactics compounded. Statistics plus fluency produced the largest joint gains.

One finding deserves more attention than it gets. The biggest beneficiaries were not the sites already ranking first. Sources sitting around position five in the underlying search results saw visibility gains above 100% after optimization, which means GEO acts as an equalizer for credible sites that never cracked the top of Google. That fluency alone moved citations by roughly 28% is equally striking: no new information, just cleaner writing, and the machine rewarded it.

Seven writing changes that follow from the data

Everything above compresses into a set of concrete habits. None of them require new tools, but all of them run against instincts built during the long-form, keyword-first era.

1.  Lead with the answer. Put the direct answer in the first 80-100 words of the page and the first sentence of every section. Chunks that bury the payoff under context lose the similarity match to chunks that state it plainly.

2.  Write self-contained passages. Aim for 50-150 word blocks where no sentence depends on the previous paragraph to make sense. ZipTie's analysis tied this format to 2.3x more citations than unstructured prose.

3.  Attach numbers to claims. "Improved conversion significantly" is invisible. "Lifted conversion 47% over six months" is extractable. Statistics addition was the single strongest tactic in the Princeton tests.

4.  Name your sources. Citing credible references inside your own content raises the odds of being cited yourself. It reads as a paradox and tested as one of the top three tactics.

5.  Make entities explicit. Fan-out queries steer toward named things: products, people, places, specs. Long pronoun chains ("it", "this approach", "the tool") break the match. Repeat the actual names.

6.  Consolidate, don't fragment. ChatGPT picks one page per domain per claim. Twenty thin pages give twenty weak candidates. Put the strongest complete answer on a single page and let it win.

7.  Date it and refresh it. Visible publish and update dates matter, especially on Perplexity, which skews heavily toward recent content. A quarterly refresh of numbers keeps a page in the recency window.

The technical layer: get retrieved before worrying about being cited

Writing changes only pay off if the engines can reach the content. A surprising number of sites fail here silently, usually through legacy robots.txt rules or firewall defaults that were never reviewed after AI crawlers appeared. The relevant user agents are worth checking individually.

CrawlerOperatorWhat blocking it costs
GPTBotOpenAIExclusion from model training data, which shapes long-term parametric knowledge of your brand
OAI-SearchBotOpenAIExclusion from ChatGPT search results and live citations
ChatGPT-UserOpenAIPages fail when a user asks ChatGPT to open your URL directly
PerplexityBotPerplexityExclusion from Perplexity's index and answer citations
ClaudeBotAnthropicExclusion from Claude's retrieval and search features

Beyond access, structure helps the chunking stage. Clean heading hierarchies, real HTML tables instead of table screenshots, Q&A formatted sections and schema markup all make passages easier to isolate and attribute. Server-side rendering matters too: content that only exists after JavaScript execution is invisible to several of these crawlers. The llms.txt standard gets discussed constantly, but adoption is early and no major engine has confirmed it as a retrieval signal, so it belongs in the "cheap bet" column rather than the strategy column.

What people tracking this are saying

"Citations are correlated, but not causal with brand appearances in the results." His broader research argues discovery now spans dozens of platforms, and cautions against treating citation counts as the whole visibility picture.

Rand Fishkin  ·  Cofounder, SparkToro

Commenting on Reddit's collapse from roughly 60% to 10% of ChatGPT citations in autumn 2025, Semrush's head of organic and AI visibility attributed the drop to ChatGPT deliberately reducing its bias toward over-cited sites, making the system harder to manipulate.

Sergei Rogulin  ·  Semrush, 3-Month Citation Study

"Including citations, quotations from relevant sources, and statistics can significantly boost source visibility." The KDD 2024 paper that coined the term GEO, and the closest thing this discipline has to a controlled experiment.

Aggarwal et al.  ·  Princeton GEO Study, KDD 2024

Solis reports finding AI crawlers blocked without teams noticing and JavaScript hiding entire site sections from retrieval. Her advice: verify actual AI visibility instead of assuming it, write in real customer language, and take community discussion seriously, because models weigh it as an authority signal.

Aleyda Solis  ·  Founder, Orainti

What reliably fails

The tracking spreadsheet accumulated a graveyard of tactics that felt productive and did nothing. The research explains why each one fails.

• Keyword stuffing. Tested directly in the Princeton study and produced no visibility gain. Semantic retrieval matches meaning, not term frequency, so the classic reflex is dead weight.

• Spreading one answer across many thin pages. Domain deduplication means the engine selects your single best page per claim. Fragmentation hands the citation to a competitor who consolidated.

• Optimizing for one engine and assuming coverage. With ~11% domain overlap between ChatGPT and Perplexity, single-platform tracking leaves most of the citation landscape unmeasured.

• Treating results as stable. BrightEdge data shows most citation positions hold week to week, but among the ones that move, the large majority are declines. Set-and-forget content decays.

• Walls of clever prose. Fluency in the GEO study meant clarity, not style. Dense, winding paragraphs actively suppress citation odds even when the underlying facts are strong.

Final Verdict

Six months of Monday-morning prompt logging leaves a clear impression: the engines are not mysterious, they are just strict. Every page that earned a citation in the tracking sheet shared the same anatomy. A direct claim, a specific number, a named source, all packed into a passage short enough to lift out whole. Every page that stayed invisible, including some that rank beautifully in Google, buried its best material under throat-clearing.

The honest takeaway is that this is not a new dark art. It is closer to old-fashioned reporting discipline enforced by software. The Princeton numbers, the Semrush volatility data and the overlap studies all reward the same behavior: say something checkable, say it early, say where it came from, and keep it current. Writers who already worked that way have been quietly winning citations without changing a thing. Everyone else has a very fixable problem, and the fix starts with the next paragraph they publish, not with a new tool subscription.

Related Posts