AI Tools

We Tested Popular AI Detectors on Human and AI-Written Content, Here's What Happened

I have a folder on my desktop called "detector jail." It holds a 2017 blog post I wrote about a week of commuter train delays, a college essay from before ChatGPT existed, and a client case study that took me two full days to research. Every file in that folder has been flagged as AI-generated by at least one detection tool. That folder is the reason this test exists.

Over nine days in late August, I fed 60 writing samples through six of the most searched AI detectors: GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, and QuillBot's free checker. Some samples were fully human, written years before generative AI was a thing. Some were raw output from ChatGPT, Claude, and Gemini. The rest were the messy middle nobody talks about, drafts started by a model and rewritten by a person. The results changed how I use these tools, and in one case, made me stop using a tool entirely.

How the test was set up

Sample groupCountWhat went in
Human20Pre-2022 blog posts, old essays, and fresh writing done in a text editor with version history on
Pure AI20Unedited output from ChatGPT, Claude, and Gemini across blog, essay, and product copy formats
Blended20AI first drafts rewritten by hand, plus human drafts lightly polished with Grammarly

Every sample ran between 500 and 900 words, because most detectors get shaky below that range. The AI samples reused briefs from a month of running identical writing briefs through different models, so the prompts reflected real assignments rather than "write me an essay about dolphins." Each sample went through every detector on the same day, and I logged the exact percentage score, not just the verdict.

One rule kept the test honest: a human sample counted as falsely flagged if the tool rated it above 50 percent AI. A soft "possibly AI" label on genuine writing still causes real damage when a client or professor is reading the report, so I did not grade on a curve.

Results at a glance

Scores show how each tool handled the three sample groups. "Caught" means the tool rated pure AI text above its own AI threshold. "False flags" counts human samples rated above 50 percent AI.

DetectorAI text caughtFalse flags on human textBlended draftsStarting price
GPTZero19 / 201 / 20Cautious, leaned humanFree tier, paid from $10/mo
Originality.ai20 / 203 / 20Aggressive, leaned AI$14.95/mo, no free tier
Copyleaks18 / 201 / 20InconsistentFrom $9.99/mo
Winston AI16 / 202 / 20Overflagged heavilyFrom $12/mo, 14-day trial
ZeroGPT17 / 205 / 20Wild score swingsFree, 15,000 characters
QuillBot16 / 200 / 20Weakest of the groupFree

Two things jumped out immediately. First, catching obvious AI text is close to a solved problem. Raw, unedited ChatGPT output got caught by every tool most of the time. Second, the real separation between these products happens on human writing and blended drafts, which is exactly where the stakes are highest.

What happened with each detector

GPTZero

Free for 10,000 words/month, paid from $10/mo

GPTZero was the closest thing to a trustworthy referee in this test. It caught 19 of 20 pure AI samples and wrongly flagged only one human piece, an oddly structured listicle I wrote in 2019. The sentence-level highlighting was genuinely useful: instead of a single scary number, it showed me which paragraphs pushed the score up, and those paragraphs usually were the most formulaic parts of the draft.

The company published a 3,000-sample benchmark in 2026 claiming 99.3 percent accuracy with a 0.24 percent false positive rate. My smaller test did not hit those numbers, but the general shape held: strong recall, rare false accusations. Worth noting for anyone tracking the industry, GPTZero was acquired by Superhuman under the Grammarly umbrella in June 2026, and the free tier still gives you 10,000 words a month.

What worked

+ Lowest false flag rate among the paid-grade tools

+ Sentence highlighting explains the score

+ Generous free tier for occasional checks

What did not

-  Missed one Claude sample written in a casual voice

-  Blended drafts sometimes sailed through untouched

Originality.ai

 $14.95/mo, credit-based, no free tier

If your only goal is making sure zero AI text slips through, Originality.ai is the strictest gatekeeper here. It caught all 20 pure AI samples, including two that fooled half the field. Independent studies have measured it around 97 percent accurate on AI detection, and my results back that up.

The cost of that strictness showed up on the human side. Three of my genuine samples got flagged, and two of them shared one trait: they had been run through Grammarly for cleanup. That matches a pattern researchers documented in 2025, where lightly polished human text caused this detector's false positive behavior to spike. If your writers use any editing assistant at all, expect friction.

What worked

+ Perfect recall on unedited AI output

+ Plagiarism and fact-check tools in one dashboard

+ Chrome extension checks drafts inside Google Docs

What did not

-  Punishes human text polished with editing tools

-  No free tier, so casual users cannot trial it properly

Copyleaks   

From $9.99/mo

Copyleaks landed in respectable middle ground. It caught 18 of 20 AI samples and flagged only one human piece, roughly matching the one-in-twenty human misclassification rate that showed up in a large third-party benchmark this year. Multilingual support is a real strength if your team publishes outside English.

Blended drafts were its weak spot. The same rewritten article scored 34 percent AI on Monday and 61 percent on Thursday after minor edits that changed fewer than 40 words. When a score moves that much on that little, I stop trusting the number and start trusting my own read of the text.

What worked

+ Cheapest entry price among the paid tools

+ Strong multilingual detection

+ Low false flag count on fully human text

What did not

-  Scores on edited drafts swung unpredictably

-  Missed two AI samples other tools caught

Winston AI

From $12/mo, 14-day trial

Winston advertises 99.98 percent accuracy on its homepage. My test, and every independent benchmark I could find, tells a humbler story. University research datasets have measured it between 76 and 83 percent, and my run landed in the same neighborhood: 16 of 20 AI samples caught, 2 human pieces wrongly flagged.

Its biggest problem was blended content. One draft that was roughly 40 percent human rewriting came back rated almost entirely AI, which would be a firing offense at some agencies if a manager took the score at face value. The OCR feature that reads scanned and handwritten documents is genuinely rare and useful for educators, but the core detector did not justify the confidence its marketing projects.

What worked

+ OCR handles scanned pages and handwriting

+ Includes plagiarism and AI image checking

+ Clean sentence-by-sentence prediction map

What did not

-  Accuracy claim far ahead of test performance

-  Treated heavily edited drafts as pure AI

ZeroGPT

Free for 15,000 characters, no signup

ZeroGPT is the tool that put my old train delay post in detector jail, and this test explained why. Five of my 20 human samples came back flagged as AI, the worst false positive performance in the group, and consistent with independent reviews that measured its false flag rate between 15 and 26 percent. One 2016 essay scored 74 percent AI. It was written six years before ChatGPT launched.

The frustrating part is that its raw AI detection is not terrible, catching 17 of 20 machine samples. But a detector that regularly accuses innocent writers is worse than no detector, because people act on those accusations. Fine for a curiosity check, unusable for any decision that affects a person.

What worked

+  Zero friction, no account needed

+  Decent recall on unedited AI text

What did not

-   Highest false flag rate in the entire test

-   Scores felt arbitrary on anything nuanced

QuillBot AI Detector

Free, capped per scan

QuillBot's free checker surprised me in one direction: it never once accused my human writing of being AI. Not a single false flag across 20 samples. Independent journalists have measured it around 80 percent overall accuracy, and the tradeoff for its caution showed on the other side of the ledger, where it missed 4 of 20 AI samples and struggled badly with blended drafts, agreeing with external tests that put its hybrid-content accuracy near 60 percent.

Credit where due, QuillBot prints a warning telling users never to rely on detection alone for decisions that affect someone's career or academics. It is the only vendor in this test that undersells its own product, and honestly, that earned some respect.

What worked

+ Only tool with zero false accusations

+ Completely free and fast

+ Honest disclaimer about its own limits

What did not

-  Weakest at catching edited AI content

-  Word cap per scan forces chunking long pieces

Where every single detector struggled

The pattern that mattered most had nothing to do with which tool won. All six stumbled on the same three types of writing, just to different degrees.

Content typeWhat the detectors didWhy it matters
AI drafts rewritten by a personVerdicts split almost randomly across tools, same text scoring 8% on one and 72% on anotherThis is how most professional content gets made now
Human text polished with editing toolsGrammarly-cleaned samples triggered flags on three detectorsWriters get punished for basic proofreading
Plain, formulaic human writingSimple sentence structures and predictable phrasing raised scores everywhereClear writing looks statistically similar to AI writing

My blended samples went through the same routine I use on paid work, drafting with a model and then rewriting the draft in structured editing passes until the voice, examples, and claims are mine. Detectors could not agree on what that output was, because honestly, there is no clean answer. It is both. The tools force a binary verdict onto writing that no longer fits a binary.

The research says the same thing, louder

My 60-sample test is small, so I checked it against the published record, and the record is blunt. A widely cited academic study from Weber-Wulff and colleagues tested 14 detection tools and found every one of them scored below 80 percent accuracy, with only five clearing 70. A Stanford study found that popular detectors flagged around 61 percent of essays written by non-native English speakers, because simpler vocabulary reads as machine-like to these models.

The institutions with the most at stake have already voted. Vanderbilt, Michigan State, Northwestern, and several other universities disabled Turnitin's AI detection over false positive concerns, even though Turnitin itself claims a false positive rate under 1 percent. When schools that pay for a tool switch it off rather than risk wrongly accusing students, that tells you more than any accuracy percentage on a sales page.

The one-line summary of every credible study: detectors are probability estimators, not lie detectors. They measure how statistically predictable your writing is, and predictable is not the same thing as artificial.

What review platforms say about these tools

Before wrapping up, I pulled the public ratings for all six detectors from Trustpilot, G2, and Capterra to see whether other users' experiences matched mine. The split between platforms turned out to be one of the most revealing parts of this whole project. Ratings below were checked in September 2026, with review counts rounded.

DetectorTrustpilotG2Capterra
GPTZero2.2 / 5  (138 reviews)4.3 / 5  (101 reviews)No rated listing
Originality.ai4.6 / 5  (1,000+ reviews)4.4 / 5  (180+ reviews)No rated listing
Copyleaks2.6 / 5  (346 reviews)4.3 / 5  (26 reviews)No rated listing
Winston AI3.9 / 5  (16 reviews)4.4 / 5  (13 reviews)No rated listing
ZeroGPT1.2 / 5  (110 reviews)Not listedNo rated listing
QuillBot4.8 / 5  (14,000+ reviews)4.4 / 5  (72 reviews)No rated listing

Three things stand out in that table. First, look at GPTZero and Copyleaks: both score well above 4 on G2, where verified business buyers leave reviews, and collapse to the low 2s on Trustpilot, where anyone can post. Reading through those Trustpilot pages, a large share of the one-star reviews come from writers and students who were falsely flagged. The people buying the tool and the people judged by it live in two different realities, and the ratings show it.

Second, ZeroGPT's 1.2 on Trustpilot is the lowest score I have ever seen for a tool this widely used, and it lines up exactly with my false flag results. Third, a caveat on QuillBot: its 4.8 covers the entire writing suite, paraphraser included, not the detector alone, so treat that number as brand sentiment rather than a detection scorecard. As for the empty Capterra column, none of these six maintains a rated Capterra profile, which tracks with how they sell: direct to individuals, educators, and content teams rather than through the software procurement channels Capterra serves.

How I actually use detection scores now

After this test, I settled on a workflow that treats detectors as smoke alarms rather than judges. Three rules cover it.

RuleIn practice
Never trust one toolRun GPTZero plus one other detector. A single verdict means little, two independent tools agreeing is worth investigating
Read the highlighted sectionsFlagged sentences are usually the flattest writing in the piece. Whether AI wrote them or not, they deserve a rewrite
Never punish on a score aloneAsk for drafts, version history, or a quick conversation about the piece before accusing anyone of anything

Final verdict after nine days of testing

Going in, I expected to crown a winner and move on. Coming out, my honest take is more uncomfortable: the best detector in this test, GPTZero, is genuinely good at one narrow job, spotting raw machine output, and even it went quiet the moment a human touched the draft. If someone pastes ChatGPT output and hits publish, these tools will catch them. If someone uses AI the way most working writers actually use it, the tools mostly shrug.

My rankings, for what they are worth after living inside these dashboards for nine days: GPTZero if you need one dependable tool, QuillBot if you need free and would rather miss AI text than accuse a real writer, Originality.ai if you run a content operation and can tolerate some false alarms as the price of strictness. I deleted my ZeroGPT bookmark on day six and have not missed it.

The bigger lesson sits outside the tools entirely. Detector jail taught me that scores punish flat, predictable writing regardless of who produced it. So the fix is the same whether a model drafted your piece or you did: add the specifics, the opinions, and the firsthand detail that no probability model expects. That kind of writing does not just pass detectors. It is the only kind worth publishing.

Related Posts