Honest second opinion — no card, no account

Is GPTZero accurate? Read its numbers right.

Paste the same text here. Compare reports, not reputations.

Try a sample:0 words · up to 2,500 per scan
Sensitive scancatches more · may flag human writing

50 free scans a day · 2,500 words per scan · markhuman is a competing detector

Reading the research

Four questions that decode any accuracy claim

These apply to GPTZero's published evaluations, to every rival's marketing page, and to anything we could publish about ourselves. A reader armed with these four cannot be misled by a decimal again.

Ask what's in the test set

GPTZero, to its credit, publishes more about its evaluation than most vendors. The questions that decide whether any published number applies to you are always the same: which models generated the AI half, how long were the documents, was anything edited or paraphrased, and whose human writing formed the control group. A number is only as transferable as its test set is similar to your text.

Separate the two error rates

A single 'accuracy' figure blends two very different failures: AI text called human, and human text called AI. For a student, the second is the one that ruins a semester; for an editor, the first is the one that matters. Any evaluation worth reading — GPTZero's or anyone's — reports the two separately, and any number that doesn't is hiding the trade-off somewhere.

Check the date

Every detector is a snapshot of the model landscape at training time. An accuracy figure measured before the current generation of models existed says little about output from them. This cuts both ways: old evaluations understate current tools on old models, and overstate them on new ones. The half-life of any published detector benchmark is months, not years.

Watch for threshold games

Detectors choose an operating point: flag more and catch more (but accuse more innocents), or flag less and stay safe (but wave more AI through). Accuracy claims are made at a chosen threshold, and the tool you use in practice may sit at a different one. When a number seems too good, the threshold is usually where the flexibility lives — for every vendor.

Held to these four questions, GPTZero comes out better than much of the market — publishing methodology at all is rare here. It does not come out exempt: its numbers remain self-measured, dated to a model landscape that moves monthly, and blended across error types whose costs fall on different people. The same sentences would be true of us, which is why we publish evidence per-document instead of a percentage per-homepage.

The shared frontier

The text that humbles every detector

Accuracy debates about individual tools miss that the hardest cases are hard for the category. If your text is on this list, weight any verdict — GPTZero's, ours, anyone's — accordingly.

01

Short passages

Under a few hundred words, the statistics that separate human from model text simply have too little to work with. Every serious tool enforces a minimum; every verdict near it is soft.

02

Mixed authorship

A human draft a model tightened, or a model draft a human rewrote, genuinely carries both signatures. Binary formats round this case away; it is the most common real document and the least certain verdict.

03

Second-language writers

Peer-reviewed research found detectors flagging non-native English writing at sharply elevated rates. The bias is a property of how detection works, and no vendor — us included — has eliminated it.

Second opinion

Compare reports, not reputations

If you just got a GPTZero score you're unsure about, run the same text in the checker at the top of this page. Agreement strengthens the verdict; disagreement tells you the text is ambiguous and deserves a human read. Our report shows its work per sentence either way.

Where we stand

A competitor wrote this page. Here are its rules

markhuman competes with GPTZero, and you should read this page knowing that. Its rules: every claim about GPTZero is public information or GPTZero's own statement, attributed as such; no head-to-head benchmark appears here because no independent study we could cite covers both tools fairly; and every limitation described — short-text weakness, threshold trade-offs, non-native false positives — applies fully to our own detector. The fifteen-minute self-test is the only accuracy evidence this page endorses.

Related reading: the ZeroGPT accuracy page applies the same rules to the other big free checker, the Turnitin explainer covers the tool most students actually face, and the GPTZero alternative page makes the switching case if that's the question you came with.

FAQ

GPTZero accuracy, answered plainly

What its published numbers mean, when it can be wrong, and what to do when detectors disagree.

Is GPTZero accurate?

On its home turf — long, unedited, English model output — GPTZero, like other serious detectors, performs well. Away from that turf (short texts, heavy edits, mixed human-AI documents, non-native English writers), its reliability drops, as does every detector's, ours included. GPTZero publishes its own accuracy research, which is more transparency than most of the field offers; those numbers are still self-measured on self-chosen test sets, and should be read with the questions this page lays out.

Can GPTZero be wrong?

Yes, in both directions, and its own team says so — GPTZero's guidance recommends against using its scores as the sole basis for disciplinary action, which is the correct advice for the entire category. False positives concentrate on polished, uniform, and second-language writing; false negatives concentrate on edited, paraphrased, and deliberately 'humanized' output. Any detector that claimed otherwise would be marketing past its own error bars.

Why did GPTZero flag my own writing as AI?

The same reasons any detector flags human work: your style is statistically smooth (formal training does this), the sample was short, English is your second language, or grammar tooling rewrote enough sentences to shift the texture. It reads style, not history. If a flag like this has real stakes — an accusation, a grade — process evidence (version history, drafts) settles what no score can.

GPTZero and another detector disagree. Which one is right?

Possibly neither, in the sense you mean. Different classifiers with different training data and thresholds disagree routinely on borderline text; the disagreement itself is the information. Text that splits two serious detectors is genuinely ambiguous — usually mixed, edited, or short — and the honest next step is a human read, not a tiebreaker vote from a third tool.

Does GPTZero still use perplexity and burstiness?

Those two metrics made GPTZero famous in early 2023, and it has said its system has evolved well beyond them — modern detection generally means trained classifiers, not two-metric rules. We describe what its current internals do no further than that, because we don't know them, and neither does anyone else outside the company. What can be said with confidence: every detector today, ours included, learns many signals rather than reading one or two.

How is markhuman different from GPTZero?

Format more than magic. markhuman commits to one of three named verdicts — Human-Written, AI-Assisted, AI-Generated — so mixed documents get named rather than scored into ambiguity, and it shades every sentence in place so you can see which lines carried the verdict. It is free without an account (50 scans/day, up to 2,500 words, 50-word minimum). The hard cases are still hard: we do not claim to escape the limits this page describes.

Evidence beats reputation

Run text you know the truth about through both tools and count the errors yourself. Free here — 2,500 words per scan, 50 scans a day.