The statistic looked perfect. I had asked a chatbot to help me draft a client post about reader trust, and it handed me a 2019 Nielsen report finding that 64 percent of readers abandon a publication after a single factual error. Clean number, believable source, exactly the point I needed. I spent twenty minutes hunting for that report before accepting the obvious: it does not exist. The model built me a fact the same way it builds a sentence, one plausible word at a time.
That was two years and a few hundred drafts ago. I still write with ChatGPT, Claude, and Gemini every working day, for client blogs, newsletters, and product pages. The tools have made me faster in ways I would not give back. They have also taught me, occasionally in front of an editor, that speed and accuracy come from two different places. The drafting can be the machine's job. The truth of the thing is still entirely mine.
What follows is the routine I actually run before anything ships. Not theory. Specific checks, in order, plus the documented cases that show what skipping them costs. Budget about fifteen minutes per article and a stubborn attitude toward primary sources.
| 01 Why models invent things | 06 Tools that help, and their limits |
| 02 What it costs when nobody checks | 07 Guardrails for teams |
| 03 The six-step verification pass | 08 The pre-publish checklist |
| 04 Where to check each kind of claim | 09 Questions I get asked |
| 05 Prompting habits that cut errors | 10 My verdict after two years |
Why models invent things
None of it makes sense until you accept what a language model does. It is not looking anything up. It predicts the next word from patterns in training data, so a citation-shaped hole gets a citation-shaped guess. The guess arrives in the same confident tone as everything else. Four habits of the technology explain most of the errors I catch.
Prediction, not retrieval Ask for a source and the model produces something that looks like one: real author, real journal, invented paper. The format is flawless because the format is what it learned. |
Trained to guess OpenAI's own researchers argued in a September 2025 paper that models bluff because standard benchmarks score a confident wrong answer above an honest "I don't know." Guessing wins points, so guessing is what you get. |
Frozen in time Every model has a knowledge cutoff. Prices, laws, job titles, and version numbers drift past it, and the model states last year's world in the present tense. |
Fluency hides failure A wrong fact and a right fact read identically. Your ear, trained by decades of well-edited prose, is the worst possible detector. The workflow below never relies on how a sentence sounds. |
What it costs when nobody checks
None of these are hypotheticals. Each row is a documented case where AI-drafted material shipped without a proper check. I keep it pinned above my desk; clients stopped arguing about verification time once they had seen it.
| YEAR | WHO | WHAT SLIPPED THROUGH | WHAT IT COST |
|---|---|---|---|
| 2023 | Two New York lawyers | A court brief citing six cases ChatGPT invented, complete with fake quotes, in Mata v. Avianca | A $5,000 sanction and a permanent spot in every AI ethics lecture |
| 2023 | CNET | Dozens of AI-written finance explainers published with light review | Corrections issued on 41 of 77 stories after outside reporters flagged basic errors |
| 2023 | Sports Illustrated | Product reviews credited to authors who did not exist, with AI-generated headshots | The content vendor was dropped and the publisher's CEO was out within weeks |
| 2024 | Air Canada | A support chatbot that invented a bereavement refund policy | A tribunal ordered the airline to pay C$812; its argument that the bot was a separate legal entity failed |
| 2025 | Deloitte Australia | A AU$440,000 government report containing fabricated academic references and a made-up quote from a federal court judgment | A partial refund of about A$97,000, a corrected report, weeks of headlines |
45% of AI assistants' answers about the news contained at least one significant issue, across roughly 3,000 responses tested by 22 public broadcasters in 2025. Sourcing was the most common failure. EBU AND BBC STUDY | 58 to 82% hallucination rate for general-purpose chatbots answering specific legal questions in Stanford's research. Purpose-built legal tools still erred 17 to 33 percent of the time. STANFORD PAPER | 1,600+ court decisions worldwide involving AI-fabricated material by mid-2026, up from roughly 120 when this public database launched in May 2025. CHARLOTIN DATABASE |

The six-step verification pass
I run this pass on every AI-assisted draft; each step catches a different failure. The first time takes half an hour. Once it becomes habit, fifteen minutes covers a standard article. I work on a printed copy or in a second pane, never inside the chat that produced the draft.
01 Mark every checkable claim Highlight anything that could be wrong: numbers, dates, names, titles, quotes, superlatives, and every sentence beginning with "studies show." The rule: highlighted means unproven. Some drafts turn mostly yellow. That tells you how much is riding on the model's memory. |
02 Chase the primary source For each highlight, find the original dataset, paper, filing, or announcement, not an article about it. A blog citing a blog citing a press release is a rumor with formatting. Paste the source URL next to the claim; you will want it for the log later. |
03 Confirm citations exist before checking what they say Fabricated references usually pair a real author and a real journal with an invented title. Search the exact title in quotes on Google Scholar, resolve any DOI at doi.org, and look on the publisher's site. No match after three minutes means it goes, whatever it claimed to prove. |
04 Re-type the numbers from the source Do not compare the draft to the source. Read the source, write the figure down fresh, then compare. Models round aggressively, swap units, and move a number from one year to another. Also check which year the number describes, which is often not the year the report came out. |
05 Treat quotes as radioactive Match every quotation word for word against a transcript, recording, or official text. Models routinely stitch a paraphrase into quote marks and attach it to a real person. If I cannot find the exact wording, it becomes an unquoted summary or it gets cut. |
06 Date-stamp anything perishable Prices, product versions, laws, executive names, rankings. Add "as of" with the month and year, and save an archived snapshot at web.archive.org when the claim matters. Half the corrections I have ever issued were facts that were true when written and false by publication. |

Where to check each kind of claim
The slowest part of verification is deciding where to look. This routing list lives in a note on my desktop and has saved me more hours than any writing tool I pay for.
| CLAIM TYPE | FIRST STOP | BACKUP |
|---|---|---|
| Statistic | The original report or dataset, opened yourself | National statistics offices such as BLS, Eurostat, or ONS |
| Study or finding | The paper via its DOI at doi.org | Google Scholar; PubMed for anything medical |
| Direct quote | The transcript, video, or official record | Wire coverage from AP or Reuters |
| Law or regulation | The statute or ruling itself on official sites | CourtListener, EUR-Lex, legislation.gov.uk |
| Company claim | Official filings and the company newsroom | SEC EDGAR or the local company registry |
| Price, spec, feature | The vendor's live page, checked today | An archived snapshot, saved for the record |
| Historical fact | A reference work plus one specialist source in agreement | University and national library databases |

Prompting habits that cut errors at the source
Verification gets lighter when the draft arrives cleaner. None of them replace checking, but they shrink the pile of highlights considerably.
Ask for the claims list The "no source" rows are your risk map before you have read a word of the prose. List every factual claim in this draft with the source you would cite. Write "no source" where you have none. |
Feed it the sources yourself Grounding the model in documents you supply is the biggest accuracy gain I have measured in my own work. Use only the attached report. If the answer is not in it, say it is not. |
Turn on search, then click the links Search-connected modes help with anything recent, but the EBU study found sourcing was the most common failure: links that do not support the sentence they sit under. A citation is a claim too. Check it. |
Invite uncertainty Imperfect, since models judge their own knowledge poorly, but it reliably surfaces the softest paragraphs first. Flag anything you are less than fully confident about. |
Review in a fresh chat A conversation that produced an error tends to defend it. I paste the draft into a clean session, or a different model entirely, and ask it to attack the facts. When they disagree, I check by hand. |
Tools that help, and what they can't do
People ask constantly whether some tool can just handle this. No. A few take real weight off, as long as you are honest about where each one stops.
| TOOL | GENUINELY GOOD FOR | WILL NOT DO |
|---|---|---|
| Search-connected assistants | Recent events, current prices, finding the primary source fast | Guaranteeing the cited page supports the claim |
| Source-grounded workspaces | Answers confined to documents you upload, NotebookLM style | Anything outside those documents |
| Citation checkers | Proving a reference exists in seconds via doi.org, Crossref, or Scholar | Telling you whether the study was any good |
| The Wayback Machine | Showing what a page said on a given date | Reflecting what it says now |
| AI-writing detectors | Guessing at style | Accuracy; a detector score says nothing about whether a sentence is true |
| Grammar and style tools | Polish and consistency | Facts. Teams forget this constantly |
Guardrails for teams
Solo writers can hold this routine in their head. The moment two or more people touch a draft, accuracy has to live in the process, because "I assumed you checked it" is how every published error I have been near happened.
Keep a verification log One shared sheet per piece: claim, source URL, who checked it, when. Five minutes of admin that turns "is this right?" into a lookup instead of an argument. |
Two sources for load-bearing numbers Any figure that lands in a headline, a chart, or a client deliverable gets confirmed in two independent places. Independent means the two are not citing each other. |
Write down your AI disclosure rule Policies differ wildly. Some newsrooms have publicly ruled out publishing machine-written prose; others allow AI drafting with human verification. What matters is that your team's line exists in writing before a controversy. |
Make corrections boring A dated corrections note at the end of a piece costs nothing and builds trust. The Deloitte episode ended with a corrected report and a refund. The version where they had dug in would have ended far worse. |

The pre-publish checklist
The whole guide, compressed to what I actually look at before hitting publish. Copy it, adapt it, tape it somewhere visible.
| ☐ Every number traced to a primary source I opened myself | ☐ Superlatives like "first" or "largest" either proven or deleted |
| ☐ Every citation resolved through a DOI, Scholar, or the publisher | ☐ The draft reviewed for facts in a fresh session or second model |
| ☐ Every quote matched word for word against the record | ☐ Every outbound link opened, and it supports the exact sentence |
| ☐ Names, titles, and affiliations confirmed current this month | ☐ Anything unverifiable rewritten as opinion or cut entirely |
| ☐ Prices, versions, and legal claims stamped with an "as of" date | ☐ The verification log updated so the next person can retrace me |

My verdict after two years
I want to be straight about where I have landed, because most writing on this topic comes from people who either fear these tools or sell them. I do neither. AI drafting has roughly doubled my output, and it has never once been the reason a piece of mine was accurate. Both are true, and they stopped feeling contradictory long ago. The mental model that works for me: treat the model like a fast, extremely well-read stranger. Wonderful company, full of ideas, and you would not repeat a single number it told you at a party without checking. Every step in this guide is that instinct, written down. What surprised me is the trade. Verification costs me about fifteen minutes per article. Drafting with AI saves me well over an hour. The math has never been close, and those fifteen minutes are what let me put my name on the result. In two years the routine has caught invented studies, a misattributed quote from a real CEO, prices two revisions out of date, and one statistic that was accurate for the UK and presented as global. Write with the machine. Publish like everything is your fault, because it is. |