Every Monday since February, the same 40 prompts have gone into ChatGPT and Perplexity from this desk. Buying questions, comparison questions, troubleshooting questions, all logged in a spreadsheet alongside every source each engine cited that week. The pattern that emerged was uncomfortable for anyone raised on classic SEO. Pages that had held Google's top three for years never appeared in the answers. Meanwhile a mid-tier vendor blog with a dated stat table got named week after week, and a two-year-old forum thread outperformed a 4,000-word pillar page on the same topic.
The spreadsheet kept pointing at one conclusion: these engines do not reward pages, they reward passages. Once that clicks, most of the old writing habits start looking like liabilities. This piece lays out what the research and six months of tracking agree on, and what it means for how content gets written now.
WHAT ACTUALLY CHANGES
• Passages beat pages. Engines retrieve whole pages but cite individual chunks, so every section has to stand on its own.
• Evidence earns citations. Princeton's GEO study measured up to 40% visibility gains from adding statistics, quotes and named sources.
• The two engines barely overlap. Only about 11% of domains cited by ChatGPT are also cited by Perplexity.
• SEO is the entry ticket, not the prize. Retrieval pools are built from search indexes, but selection happens on structure and specificity.
• Volatility is the norm. Reddit went from roughly 60% of ChatGPT responses to 10% in about six weeks. Nothing here is set-and-forget.
Four numbers that explain the whole game
The AI citation research space matured fast between 2024 and 2026. Several large, independent datasets now exist, and they converge on the same story: visibility inside AI answers follows its own rules, and those rules are measurable.
| Number | What it measures | Source |
| 40% | Maximum visibility lift from GEO tactics like adding statistics and cited sources | Princeton, KDD 2024 |
| 11% | Domain overlap between ChatGPT and Perplexity citations across 680M analyzed citations | Averi, 2026 |
| 21.9 | Average citations per Perplexity answer, versus roughly 8.3 for ChatGPT | Whitehat SEO, 2026 |
| 2.3x | More citations for self-contained 50-150 word chunks than unstructured text | ZipTie.dev, 2026 |

How ChatGPT and Perplexity actually pick sources
Neither engine reads a prompt and searches for it verbatim. When ChatGPT decides a question needs live information, it breaks the prompt into multiple sub-queries, a process the industry calls query fan-out, and sends them to Bing and its own retrieval systems (Ahrefs, 2026). Perplexity runs a comparable pipeline against its own hybrid index and cites far more densely, attaching sources to individual claims rather than to the answer as a whole.
What happens next is the part most writers underestimate. The engines do not evaluate whole pages. Retrieved pages get split into chunks, chunks get converted into vectors, and the passages most semantically similar to each sub-query get pulled into the answer. Testing by SEO researcher Metehan Yeşilyurt puts ChatGPT's retrieval window at roughly 38 to 65 sources per search, and the vast majority of retrieved pages never earn a citation because the final selection binds to specific sentences (SERPsecrets, 2026).
1. Query fan-out. One prompt becomes several sub-queries. A page ranking for "best CRM" can lose to a page matching the sub-query "CRM for two-person real estate teams".
2. Retrieval. Search indexes assemble a candidate pool. Classic SEO signals still gate this stage: unindexed or blocked pages score zero.
3. Chunking and selection. Pages get split into passages and compared by semantic similarity. Claim density, entity clarity and recency decide which chunks survive.
4. Attribution. The winning passages get cited inline. One important quirk: ChatGPT deduplicates by domain, so it picks a single best page per site per claim.
The two engines apply this pipeline to very different source pools, which is why optimizing for one does not automatically cover the other.

| Signal | ChatGPT | Perplexity |
| Retrieval backbone | Bing plus proprietary systems, with query fan-out into sub-queries | Own hybrid index with real-time retrieval |
| Citations per answer | ~8.3 on average | ~21.9 on average, the densest of the major engines |
| Favored sources | Wikipedia and encyclopedic references, LinkedIn (about 14.3% of responses per Semrush's 325K-prompt study), institutional sites | Reddit and community discussion (roughly 20% concentration per Evertune), tier-one journalism, LinkedIn at 5.3% |
| Freshness sensitivity | Moderate; leans on established authority | High; one 2026 analysis found strong citation rates for content under 30 days old |
| Overlap with each other | Only ~11-12% of cited domains appear on both engines | Per Averi's 680M-citation analysis and Passionfruit's cross-platform check |
What the Princeton GEO study actually proved
The academic anchor for all of this is "GEO: Generative Engine Optimization" by Aggarwal and colleagues, presented at KDD 2024 and available as arXiv preprint 2311.09735. The team built a benchmark of 10,000 queries across multiple domains, tested nine content modifications on a system built to mimic Bing Chat, and validated the strongest tactics on Perplexity.
The results split cleanly into tactics that moved citations and tactics that did nothing or backfired.
| Tactic | What it means in practice | Measured effect |
| Statistics addition | Replacing vague claims with specific, sourced numbers | up to +40% |
| Quotation addition | Adding attributed quotes from credible voices | +30 to 40% |
| Cite sources | Linking claims to named external references | ~+28% and up |
| Fluency optimization | Rewriting for clearer, easier-to-parse prose, no new facts added | ~+28% |
| Authoritative voice | Confident, direct phrasing instead of hedging | positive |
| Keyword stuffing | Adding more target keywords, the classic SEO reflex | no gain |
Combined tactics compounded. Statistics plus fluency produced the largest joint gains.
One finding deserves more attention than it gets. The biggest beneficiaries were not the sites already ranking first. Sources sitting around position five in the underlying search results saw visibility gains above 100% after optimization, which means GEO acts as an equalizer for credible sites that never cracked the top of Google. That fluency alone moved citations by roughly 28% is equally striking: no new information, just cleaner writing, and the machine rewarded it.
Seven writing changes that follow from the data
Everything above compresses into a set of concrete habits. None of them require new tools, but all of them run against instincts built during the long-form, keyword-first era.

1. Lead with the answer. Put the direct answer in the first 80-100 words of the page and the first sentence of every section. Chunks that bury the payoff under context lose the similarity match to chunks that state it plainly.
2. Write self-contained passages. Aim for 50-150 word blocks where no sentence depends on the previous paragraph to make sense. ZipTie's analysis tied this format to 2.3x more citations than unstructured prose.
3. Attach numbers to claims. "Improved conversion significantly" is invisible. "Lifted conversion 47% over six months" is extractable. Statistics addition was the single strongest tactic in the Princeton tests.
4. Name your sources. Citing credible references inside your own content raises the odds of being cited yourself. It reads as a paradox and tested as one of the top three tactics.
5. Make entities explicit. Fan-out queries steer toward named things: products, people, places, specs. Long pronoun chains ("it", "this approach", "the tool") break the match. Repeat the actual names.
6. Consolidate, don't fragment. ChatGPT picks one page per domain per claim. Twenty thin pages give twenty weak candidates. Put the strongest complete answer on a single page and let it win.
7. Date it and refresh it. Visible publish and update dates matter, especially on Perplexity, which skews heavily toward recent content. A quarterly refresh of numbers keeps a page in the recency window.
The technical layer: get retrieved before worrying about being cited
Writing changes only pay off if the engines can reach the content. A surprising number of sites fail here silently, usually through legacy robots.txt rules or firewall defaults that were never reviewed after AI crawlers appeared. The relevant user agents are worth checking individually.
| Crawler | Operator | What blocking it costs |
| GPTBot | OpenAI | Exclusion from model training data, which shapes long-term parametric knowledge of your brand |
| OAI-SearchBot | OpenAI | Exclusion from ChatGPT search results and live citations |
| ChatGPT-User | OpenAI | Pages fail when a user asks ChatGPT to open your URL directly |
| PerplexityBot | Perplexity | Exclusion from Perplexity's index and answer citations |
| ClaudeBot | Anthropic | Exclusion from Claude's retrieval and search features |

Beyond access, structure helps the chunking stage. Clean heading hierarchies, real HTML tables instead of table screenshots, Q&A formatted sections and schema markup all make passages easier to isolate and attribute. Server-side rendering matters too: content that only exists after JavaScript execution is invisible to several of these crawlers. The llms.txt standard gets discussed constantly, but adoption is early and no major engine has confirmed it as a retrieval signal, so it belongs in the "cheap bet" column rather than the strategy column.
What people tracking this are saying
"Citations are correlated, but not causal with brand appearances in the results." His broader research argues discovery now spans dozens of platforms, and cautions against treating citation counts as the whole visibility picture. Rand Fishkin · Cofounder, SparkToro |
Commenting on Reddit's collapse from roughly 60% to 10% of ChatGPT citations in autumn 2025, Semrush's head of organic and AI visibility attributed the drop to ChatGPT deliberately reducing its bias toward over-cited sites, making the system harder to manipulate. Sergei Rogulin · Semrush, 3-Month Citation Study |
"Including citations, quotations from relevant sources, and statistics can significantly boost source visibility." The KDD 2024 paper that coined the term GEO, and the closest thing this discipline has to a controlled experiment. Aggarwal et al. · Princeton GEO Study, KDD 2024 |
Solis reports finding AI crawlers blocked without teams noticing and JavaScript hiding entire site sections from retrieval. Her advice: verify actual AI visibility instead of assuming it, write in real customer language, and take community discussion seriously, because models weigh it as an authority signal. Aleyda Solis · Founder, Orainti |
What reliably fails
The tracking spreadsheet accumulated a graveyard of tactics that felt productive and did nothing. The research explains why each one fails.
• Keyword stuffing. Tested directly in the Princeton study and produced no visibility gain. Semantic retrieval matches meaning, not term frequency, so the classic reflex is dead weight.
• Spreading one answer across many thin pages. Domain deduplication means the engine selects your single best page per claim. Fragmentation hands the citation to a competitor who consolidated.
• Optimizing for one engine and assuming coverage. With ~11% domain overlap between ChatGPT and Perplexity, single-platform tracking leaves most of the citation landscape unmeasured.
• Treating results as stable. BrightEdge data shows most citation positions hold week to week, but among the ones that move, the large majority are declines. Set-and-forget content decays.
• Walls of clever prose. Fluency in the GEO study meant clarity, not style. Dense, winding paragraphs actively suppress citation odds even when the underlying facts are strong.
Final Verdict
Six months of Monday-morning prompt logging leaves a clear impression: the engines are not mysterious, they are just strict. Every page that earned a citation in the tracking sheet shared the same anatomy. A direct claim, a specific number, a named source, all packed into a passage short enough to lift out whole. Every page that stayed invisible, including some that rank beautifully in Google, buried its best material under throat-clearing.
The honest takeaway is that this is not a new dark art. It is closer to old-fashioned reporting discipline enforced by software. The Princeton numbers, the Semrush volatility data and the overlap studies all reward the same behavior: say something checkable, say it early, say where it came from, and keep it current. Writers who already worked that way have been quietly winning citations without changing a thing. Everyone else has a very fixable problem, and the fix starts with the next paragraph they publish, not with a new tool subscription.