Marketing

GEO explained: how to get your blog cited by ChatGPT and Perplexity

For the past five months I have started most workdays the same way: coffee, then the same forty questions typed into ChatGPT and Perplexity one by one, with a tracking sheet open on the second monitor. I log which sites each engine cites, what gets swapped week to week, and how long a fresh article takes to appear as a source. My own blog sits in that sheet too, which is how I learned the hard way that a post can rank on page one of Google and still never show up in a single AI answer.

That gap has a name now: Generative Engine Optimization, or GEO. It is not a rebrand of SEO and it is not a trick. It is a short list of specific, testable habits that make your content easier for an answer engine to retrieve, trust, and quote. Everything below comes from peer-reviewed research, platform documentation, and what my own sheet keeps confirming.

QUICK ANSWER

GEO is the practice of optimizing content to be retrieved and cited inside AI-generated answers on ChatGPT, Perplexity, and Google's AI Mode. Where SEO competes for a ranking that earns a click, GEO competes for one of the three to five citation slots inside the answer itself. The term comes from a 2024 paper by researchers at Princeton, Georgia Tech, IIT Delhi, and the Allen Institute for AI.

The two disciplines overlap but reward different things. The cleanest separation I have found:

DimensionTraditional SEOGEO
You win whenYour page ranks high and earns the clickYour page is quoted and linked inside the answer
Unit of competitionThe whole pageIndividual passages the model can extract
Core signalsLinks, keywords, engagement, technical healthFact density, source attribution, extractability, freshness
Result stabilityRankings shift graduallyCitations can change on every single query run
RelationshipComplementary. Strong SEO feeds the retrieval pool that GEO competes inside

The numbers behind the shift

AI referral traffic is still small, but the growth curve and conversion quality earn it a slice of your content budget now, not later.

StatWhat it measuresSource
16xGrowth in website traffic from AI search engines between 2024 and 2026SE Ranking, 2026
~75%Share of all AI referral traffic that comes from ChatGPT aloneSE Ranking, 2026
+31%Better conversion for AI referrals vs non-AI traffic in the 2025 holiday seasonAdobe Digital Insights
900MWeekly active ChatGPT users reported as of February 2026OpenAI

One more detail reframed my planning: Adobe measured 1.13 billion AI referral visits in a single month of 2025, up 357 percent year over year, and those visitors behaved like warm leads. Small channel, unusually high intent.

ENGINE 1 OF 2

How ChatGPT decides which sources to cite

ChatGPT runs in two modes, and only one can cite you. The default mode answers from training data with no live citations. Search mode queries the web, reads candidate pages, and attaches source links. Every citation you have seen in ChatGPT came from that second mode, through a four-stage pipeline.

ChatGPT can only cite you in search mode, when it reads the live web. Photo: Jernej Furman, CC BY 2.0, via Wikimedia Commons.

STAGE 01   Your prompt becomes many queries

ChatGPT breaks one question into several sub-queries, which practitioners call query fan-out. One prompt can pull candidate URLs from dozens of searches, so pages covering adjacent subtopics get multiple entry points.

STAGE 02   Bing supplies the candidates

OpenAI's search runs on Microsoft's index. Seer Interactive matched thousands of SearchGPT citations against Bing and found 87 percent aligned with its top listings. Not indexed in Bing means invisible here.

STAGE 03   Most candidates get cut

Retrieval is not citation. AirOps found roughly 85 percent of retrieved pages never appear in the final answer. Vague passages, thin claims, and hard-to-parse markup fail here.

STAGE 04   Passages earn the slots

The model cites sources whose sentences directly support the answer it is composing. Typical responses carry three to six citations, so self-contained, quotable passages beat long buildups.

One correlation worth knowing: AirOps found pages ranking first on Google get cited about 3.5 times more often. Google rank is not a direct input here, but the qualities that earn it survive ChatGPT's filters too. The direct input is Bing, and Bing Webmaster Tools is the single most neglected dashboard in GEO. I registered my site there in March and watched pages Google had indexed for a year finally become citable weeks later.

ENGINE 2 OF 2

How Perplexity decides which sources to cite

Perplexity is a different machine. It searches the live web on every query using its own index, built by PerplexityBot and search partners, then reranks candidates before a model writes the cited answer. Its head of search, Alexandr Yarats, has described the design goal as optimizing answers for helpfulness and factuality rather than click probability, on a compact index focused on high-quality pages.

A Perplexity answer with its Sources row and inline numbered citations, the slots GEO competes for. Screenshot: public domain, via Wikimedia Commons.

Third-party reverse-engineering studies, unconfirmed by Perplexity, describe the same funnel: five to ten candidate pages retrieved per query, three to five cited. Four signals separate the cited from the ignored:

■  Freshness, weighted hard. Analyses report a citation boost for pages updated within roughly 30 days, compressing to a couple of days on fast-moving topics. In my sheet, Perplexity swaps citations after updates far faster than ChatGPT.

■  Extraction safety. Perplexity cites what it can quote without distorting the meaning. Direct claims with numbers survive. Hedged, meandering paragraphs get skipped even on strong domains.

■  Third-party trust for commercial queries. On buying-intent questions, review platforms like G2, Capterra, and Trustpilot show up constantly. Your own product page rarely wins that slot.

■  Entity clarity. Pages that state plainly who wrote them, what the site covers, and where each claim comes from get categorized and retrieved more reliably.

ChatGPT vs Perplexity, side by side

FactorChatGPT searchPerplexity
Candidate indexBing's index via Microsoft partnershipOwn index (PerplexityBot plus partners)
Retrieval triggerOnly when search mode firesLive web search on effectively every query
Citations per answerUsually 3 to 6Usually 3 to 5
Freshness weightModerate; index lag of weeks is commonHeavy; recent updates get a visible boost
First move to makeVerify Bing indexing, fix coverage gapsAllow PerplexityBot, refresh key pages
Referral quality noteLargest AI referral source by volumeAbout 13 pages browsed per referral session, above Google's average (DOC Digital, 2026)

What the Princeton GEO study actually proved

Most GEO advice online is recycled opinion. The exception is the paper that named the field: GEO: Generative Engine Optimization, published at KDD 2024 by Aggarwal and colleagues. The team built GEO-bench, a benchmark of 10,000 queries across multiple domains, then tested nine content modifications for their effect on a source's visibility inside generated answers, validating results on Perplexity.

TacticWhat it means in practiceResult
Add statisticsReplace vague claims with specific, sourced numbersTop tier: up to 40% visibility lift
Add quotationsInclude attributed quotes from credible people or documentsTop tier: 30 to 40% lift
Cite sourcesLink and name where your own claims come fromTop tier: 30 to 40% lift
Fluency optimizationRewrite for clear, direct, readable sentencesMeaningful lift, strongest in combination
Authoritative voiceState claims confidently instead of hedgingModerate lift
Keyword stuffingCramming extra query terms into the textDid not help, sometimes hurt

That table explains a pattern I kept seeing before I understood it. The tactics that win citations are the ones that give a model verifiable material to quote: numbers, named sources, attributed statements. The one classic SEO reflex tested, keyword stuffing, was the clear loser. Effect sizes varied by domain, so treat these as strong priors to test on your own topics, not laws.

Two later findings sharpen the picture. HubSpot's AI Search Trends research found content published within the previous three months roughly three times more likely to earn citations than older pages on the same subject. And C-SEO Bench, a 2025 academic benchmark, found most tricks aimed at manipulating the model do nothing while plain source relevance keeps working. Substance compounds. Tricks decay.

The 7-move GEO playbook

The exact sequence I run on every post I want cited, technical prerequisites first, editorial work second.

01  Get properly indexed in Bing

Register in Bing Webmaster Tools, submit your sitemap, and turn on IndexNow so new URLs get pushed instantly. Then check the coverage report for pages Google has but Bing lacks.

02  Open robots.txt to answer-engine crawlers

Allow OAI-SearchBot and PerplexityBot at minimum. Many sites blanket-blocked every AI user agent in 2023 and forgot, which silently removes them from answers.

WHY: BLOCKED CRAWLER MEANS ZERO RETRIEVAL, FULL STOP

03  Lead every section with the answer

Put a two-to-three sentence direct answer immediately under each question-style H2, then elaborate. Models extract self-contained passages, and a paragraph that only makes sense after the three above it cannot be safely quoted.

WHY: PASSAGE-LEVEL EXTRACTABILITY DECIDES STAGE-THREE SURVIVAL

04  Load pages with sourced numbers and quotes

Swap “many businesses are adopting AI” for a specific figure with a named source, and add one or two attributed expert quotes. No editorial change earned more in the Princeton tests.

WHY: TOP-TIER TACTICS IN THE KDD 2024 STUDY, UP TO 40% LIFT

05  Publish something original

A small survey, a benchmark, a scraped dataset, a documented experiment. When an engine needs a figure, it prefers the primary source over the seventh summary of it. Original data is a moat summarizers cannot cross.

WHY: ORIGINATORS CONSISTENTLY OUTCITE AGGREGATORS IN CITATION DATASETS

06  Refresh on a real schedule

Update key pages with new data, new sources, and corrected claims every quarter; touch fast-moving pages monthly. Changing the date without the substance does not fool the recency signal.

WHY: 3X CITATION LIKELIHOOD FOR CONTENT UNDER 3 MONTHS OLD (HUBSPOT)

07  Earn mentions beyond your own domain

Citations concentrate on entities the engines already trust. Get your data referenced by industry publications, keep review-platform profiles accurate if you sell anything, and make your name and topic co-occur across the web.

WHY: COMMERCIAL ANSWERS LEAN ON THIRD-PARTY TRUST SURFACES

The 5-minute crawler audit

Open your robots.txt, then check it against this table. Training and search crawlers are separate, and confusing them is the most common self-inflicted GEO wound I see.

User agentWhat it doesBlock it and you lose
BingbotBuilds the Bing index that feeds ChatGPT search candidatesAll ChatGPT citation potential
OAI-SearchBotOpenAI's crawler for surfacing and linking sites in ChatGPT searchChatGPT search visibility
ChatGPT-UserFetches a page on demand when a user's request needs itLive page reads during chats
PerplexityBotBuilds Perplexity's search indexPerplexity citation potential
GPTBotCollects data for training OpenAI models; separate from searchNothing in live answers; a policy choice

Common mistakes worth naming

DO THISSKIP THIS

■  Serve your main content as server-rendered HTML that loads without JavaScript

■  Write H2s as the literal questions people ask

■  Add Article and FAQ schema plus a real author bio with credentials

■  Keep a changelog line on updated pages stating what changed and when

■  Keyword stuffing, which failed outright in the controlled study

■  Treating llms.txt as a ranking lever before any engine confirms using it

■  Paywalling or gating the exact passages you want quoted

■  Faking dateModified while leaving the content untouched

How to know it is working

Build a fixed set of 30 to 50 prompts your audience genuinely asks, run them weekly in both engines, and log whether you are cited, which page earned it, and who took the other slots. Citations are volatile between runs, so judge trends over 30 to 90 day windows. In analytics, segment chatgpt.com and perplexity.ai referrers separately, and expect clicks to understate influence, since plenty of readers see the citation and never click.

Final Verdict

The honest version: GEO rewarded slower, better work, and punished everything else I tried. The first thing that ever got my blog cited was not a hack. It was rebuilding an old comparison post around a small dataset I collected myself, every claim numbered and sourced, then submitting the URL through IndexNow. Perplexity picked it up within days. ChatGPT took closer to a month, which matched the Bing indexing lag I had read about but only believed once I watched it in my own sheet.

The shortcuts went nowhere. A control page I stuffed with query variations never earned a citation, and neither did anything hidden behind a JavaScript render the crawlers could not parse. The pattern in my data is boring and repeatable: be the source of a fact, make that fact effortless to extract, and stay reachable by the right crawlers. That is the whole discipline.

Start this month even though the traffic is small today. The engines are compounding user habits, citation authority accrues to whoever shows up early, and every study below points the same direction. I would rather own three citation slots in a growing channel than a tenth blue link in a shrinking one.

Related Posts