For the past five months I have started most workdays the same way: coffee, then the same forty questions typed into ChatGPT and Perplexity one by one, with a tracking sheet open on the second monitor. I log which sites each engine cites, what gets swapped week to week, and how long a fresh article takes to appear as a source. My own blog sits in that sheet too, which is how I learned the hard way that a post can rank on page one of Google and still never show up in a single AI answer.
That gap has a name now: Generative Engine Optimization, or GEO. It is not a rebrand of SEO and it is not a trick. It is a short list of specific, testable habits that make your content easier for an answer engine to retrieve, trust, and quote. Everything below comes from peer-reviewed research, platform documentation, and what my own sheet keeps confirming.
QUICK ANSWER GEO is the practice of optimizing content to be retrieved and cited inside AI-generated answers on ChatGPT, Perplexity, and Google's AI Mode. Where SEO competes for a ranking that earns a click, GEO competes for one of the three to five citation slots inside the answer itself. The term comes from a 2024 paper by researchers at Princeton, Georgia Tech, IIT Delhi, and the Allen Institute for AI. |
The two disciplines overlap but reward different things. The cleanest separation I have found:
| Dimension | Traditional SEO | GEO |
|---|---|---|
| You win when | Your page ranks high and earns the click | Your page is quoted and linked inside the answer |
| Unit of competition | The whole page | Individual passages the model can extract |
| Core signals | Links, keywords, engagement, technical health | Fact density, source attribution, extractability, freshness |
| Result stability | Rankings shift gradually | Citations can change on every single query run |
| Relationship | Complementary. Strong SEO feeds the retrieval pool that GEO competes inside | |
The numbers behind the shift
AI referral traffic is still small, but the growth curve and conversion quality earn it a slice of your content budget now, not later.
| Stat | What it measures | Source |
|---|---|---|
| 16x | Growth in website traffic from AI search engines between 2024 and 2026 | SE Ranking, 2026 |
| ~75% | Share of all AI referral traffic that comes from ChatGPT alone | SE Ranking, 2026 |
| +31% | Better conversion for AI referrals vs non-AI traffic in the 2025 holiday season | Adobe Digital Insights |
| 900M | Weekly active ChatGPT users reported as of February 2026 | OpenAI |
One more detail reframed my planning: Adobe measured 1.13 billion AI referral visits in a single month of 2025, up 357 percent year over year, and those visitors behaved like warm leads. Small channel, unusually high intent.
ENGINE 1 OF 2
How ChatGPT decides which sources to cite
ChatGPT runs in two modes, and only one can cite you. The default mode answers from training data with no live citations. Search mode queries the web, reads candidate pages, and attaches source links. Every citation you have seen in ChatGPT came from that second mode, through a four-stage pipeline.

ChatGPT can only cite you in search mode, when it reads the live web. Photo: Jernej Furman, CC BY 2.0, via Wikimedia Commons.
STAGE 01 Your prompt becomes many queries
ChatGPT breaks one question into several sub-queries, which practitioners call query fan-out. One prompt can pull candidate URLs from dozens of searches, so pages covering adjacent subtopics get multiple entry points.
STAGE 02 Bing supplies the candidates
OpenAI's search runs on Microsoft's index. Seer Interactive matched thousands of SearchGPT citations against Bing and found 87 percent aligned with its top listings. Not indexed in Bing means invisible here.
STAGE 03 Most candidates get cut
Retrieval is not citation. AirOps found roughly 85 percent of retrieved pages never appear in the final answer. Vague passages, thin claims, and hard-to-parse markup fail here.
STAGE 04 Passages earn the slots
The model cites sources whose sentences directly support the answer it is composing. Typical responses carry three to six citations, so self-contained, quotable passages beat long buildups.
One correlation worth knowing: AirOps found pages ranking first on Google get cited about 3.5 times more often. Google rank is not a direct input here, but the qualities that earn it survive ChatGPT's filters too. The direct input is Bing, and Bing Webmaster Tools is the single most neglected dashboard in GEO. I registered my site there in March and watched pages Google had indexed for a year finally become citable weeks later.
ENGINE 2 OF 2
How Perplexity decides which sources to cite
Perplexity is a different machine. It searches the live web on every query using its own index, built by PerplexityBot and search partners, then reranks candidates before a model writes the cited answer. Its head of search, Alexandr Yarats, has described the design goal as optimizing answers for helpfulness and factuality rather than click probability, on a compact index focused on high-quality pages.

A Perplexity answer with its Sources row and inline numbered citations, the slots GEO competes for. Screenshot: public domain, via Wikimedia Commons.
Third-party reverse-engineering studies, unconfirmed by Perplexity, describe the same funnel: five to ten candidate pages retrieved per query, three to five cited. Four signals separate the cited from the ignored:
■ Freshness, weighted hard. Analyses report a citation boost for pages updated within roughly 30 days, compressing to a couple of days on fast-moving topics. In my sheet, Perplexity swaps citations after updates far faster than ChatGPT.
■ Extraction safety. Perplexity cites what it can quote without distorting the meaning. Direct claims with numbers survive. Hedged, meandering paragraphs get skipped even on strong domains.
■ Third-party trust for commercial queries. On buying-intent questions, review platforms like G2, Capterra, and Trustpilot show up constantly. Your own product page rarely wins that slot.
■ Entity clarity. Pages that state plainly who wrote them, what the site covers, and where each claim comes from get categorized and retrieved more reliably.
ChatGPT vs Perplexity, side by side
| Factor | ChatGPT search | Perplexity |
|---|---|---|
| Candidate index | Bing's index via Microsoft partnership | Own index (PerplexityBot plus partners) |
| Retrieval trigger | Only when search mode fires | Live web search on effectively every query |
| Citations per answer | Usually 3 to 6 | Usually 3 to 5 |
| Freshness weight | Moderate; index lag of weeks is common | Heavy; recent updates get a visible boost |
| First move to make | Verify Bing indexing, fix coverage gaps | Allow PerplexityBot, refresh key pages |
| Referral quality note | Largest AI referral source by volume | About 13 pages browsed per referral session, above Google's average (DOC Digital, 2026) |
What the Princeton GEO study actually proved
Most GEO advice online is recycled opinion. The exception is the paper that named the field: GEO: Generative Engine Optimization, published at KDD 2024 by Aggarwal and colleagues. The team built GEO-bench, a benchmark of 10,000 queries across multiple domains, then tested nine content modifications for their effect on a source's visibility inside generated answers, validating results on Perplexity.
| Tactic | What it means in practice | Result |
|---|---|---|
| Add statistics | Replace vague claims with specific, sourced numbers | Top tier: up to 40% visibility lift |
| Add quotations | Include attributed quotes from credible people or documents | Top tier: 30 to 40% lift |
| Cite sources | Link and name where your own claims come from | Top tier: 30 to 40% lift |
| Fluency optimization | Rewrite for clear, direct, readable sentences | Meaningful lift, strongest in combination |
| Authoritative voice | State claims confidently instead of hedging | Moderate lift |
| Keyword stuffing | Cramming extra query terms into the text | Did not help, sometimes hurt |
That table explains a pattern I kept seeing before I understood it. The tactics that win citations are the ones that give a model verifiable material to quote: numbers, named sources, attributed statements. The one classic SEO reflex tested, keyword stuffing, was the clear loser. Effect sizes varied by domain, so treat these as strong priors to test on your own topics, not laws.
Two later findings sharpen the picture. HubSpot's AI Search Trends research found content published within the previous three months roughly three times more likely to earn citations than older pages on the same subject. And C-SEO Bench, a 2025 academic benchmark, found most tricks aimed at manipulating the model do nothing while plain source relevance keeps working. Substance compounds. Tricks decay.
The 7-move GEO playbook
The exact sequence I run on every post I want cited, technical prerequisites first, editorial work second.
01 Get properly indexed in Bing
Register in Bing Webmaster Tools, submit your sitemap, and turn on IndexNow so new URLs get pushed instantly. Then check the coverage report for pages Google has but Bing lacks.

02 Open robots.txt to answer-engine crawlers
Allow OAI-SearchBot and PerplexityBot at minimum. Many sites blanket-blocked every AI user agent in 2023 and forgot, which silently removes them from answers.
WHY: BLOCKED CRAWLER MEANS ZERO RETRIEVAL, FULL STOP
03 Lead every section with the answer
Put a two-to-three sentence direct answer immediately under each question-style H2, then elaborate. Models extract self-contained passages, and a paragraph that only makes sense after the three above it cannot be safely quoted.
WHY: PASSAGE-LEVEL EXTRACTABILITY DECIDES STAGE-THREE SURVIVAL
04 Load pages with sourced numbers and quotes
Swap “many businesses are adopting AI” for a specific figure with a named source, and add one or two attributed expert quotes. No editorial change earned more in the Princeton tests.
WHY: TOP-TIER TACTICS IN THE KDD 2024 STUDY, UP TO 40% LIFT
05 Publish something original
A small survey, a benchmark, a scraped dataset, a documented experiment. When an engine needs a figure, it prefers the primary source over the seventh summary of it. Original data is a moat summarizers cannot cross.
WHY: ORIGINATORS CONSISTENTLY OUTCITE AGGREGATORS IN CITATION DATASETS
06 Refresh on a real schedule
Update key pages with new data, new sources, and corrected claims every quarter; touch fast-moving pages monthly. Changing the date without the substance does not fool the recency signal.
WHY: 3X CITATION LIKELIHOOD FOR CONTENT UNDER 3 MONTHS OLD (HUBSPOT)
07 Earn mentions beyond your own domain
Citations concentrate on entities the engines already trust. Get your data referenced by industry publications, keep review-platform profiles accurate if you sell anything, and make your name and topic co-occur across the web.
WHY: COMMERCIAL ANSWERS LEAN ON THIRD-PARTY TRUST SURFACES
The 5-minute crawler audit
Open your robots.txt, then check it against this table. Training and search crawlers are separate, and confusing them is the most common self-inflicted GEO wound I see.
| User agent | What it does | Block it and you lose |
|---|---|---|
| Bingbot | Builds the Bing index that feeds ChatGPT search candidates | All ChatGPT citation potential |
| OAI-SearchBot | OpenAI's crawler for surfacing and linking sites in ChatGPT search | ChatGPT search visibility |
| ChatGPT-User | Fetches a page on demand when a user's request needs it | Live page reads during chats |
| PerplexityBot | Builds Perplexity's search index | Perplexity citation potential |
| GPTBot | Collects data for training OpenAI models; separate from search | Nothing in live answers; a policy choice |
Common mistakes worth naming
| DO THIS | SKIP THIS |
■ Serve your main content as server-rendered HTML that loads without JavaScript ■ Write H2s as the literal questions people ask ■ Add Article and FAQ schema plus a real author bio with credentials ■ Keep a changelog line on updated pages stating what changed and when | ■ Keyword stuffing, which failed outright in the controlled study ■ Treating llms.txt as a ranking lever before any engine confirms using it ■ Paywalling or gating the exact passages you want quoted ■ Faking dateModified while leaving the content untouched |
How to know it is working
Build a fixed set of 30 to 50 prompts your audience genuinely asks, run them weekly in both engines, and log whether you are cited, which page earned it, and who took the other slots. Citations are volatile between runs, so judge trends over 30 to 90 day windows. In analytics, segment chatgpt.com and perplexity.ai referrers separately, and expect clicks to understate influence, since plenty of readers see the citation and never click.

Final Verdict
The honest version: GEO rewarded slower, better work, and punished everything else I tried. The first thing that ever got my blog cited was not a hack. It was rebuilding an old comparison post around a small dataset I collected myself, every claim numbered and sourced, then submitting the URL through IndexNow. Perplexity picked it up within days. ChatGPT took closer to a month, which matched the Bing indexing lag I had read about but only believed once I watched it in my own sheet. The shortcuts went nowhere. A control page I stuffed with query variations never earned a citation, and neither did anything hidden behind a JavaScript render the crawlers could not parse. The pattern in my data is boring and repeatable: be the source of a fact, make that fact effortless to extract, and stay reachable by the right crawlers. That is the whole discipline. Start this month even though the traffic is small today. The engines are compounding user habits, citation authority accrues to whoever shows up early, and every study below points the same direction. I would rather own three citation slots in a growing channel than a tenth blue link in a shrinking one. |