A Winston AI alternative that shows evidence, not a decimal.
Three named verdicts and a score on every sentence — not a homepage accuracy claim.
50 free scans a day · 2,500 words per scan · text is scored in memory and never stored
Why this page exists
A decimal on a landing page is not a report
Winston AI leads with an accuracy figure — 99.98%, by its own measurement. That is a real number on a lab set, and any number we produced for ourselves would have the same shape: it describes the vendor's test, not the document in your clipboard. markhuman leads with the report instead: three named verdicts and a score on every sentence, 50 scans a day at 2,500 words, no account.
We refuse to score fewer than 50 words. Below that length, any detector's output is closer to noise than to evidence. A four-nines claim cannot transfer to a fifty-word email, a non-native draft, or a document a person reworked after a model pass — and those are the cases people actually paste.
Will this agree with Winston on the same text? Sometimes. Detectors are different classifiers trained on different corpora with thresholds drawn in different places. Disagreement is normal. Where two tools split, the honest reading is that the text is genuinely ambiguous and deserves a human look at the sentences that carried the signal — not that one vendor is lying.
We do not publish head-to-head benchmarks against Winston or anyone else, and this page contains no accuracy percentage of our own. Vendor-measured numbers describe the vendor's own test set, not your document. Short passages, non-native English, and heavily edited drafts are hard cases for the entire field, ours included. Text is scored in memory and never stored.
The four-nines problem
What an accuracy claim
can and cannot promise
This is not an accusation — Winston's figure is presumably a real measurement on its own test set, and any number we produced for ourselves would have the same shape. The problem is structural: a lab accuracy figure, anyone's, cannot transfer to the one document you are about to paste.
The number describes the test set
An accuracy figure is measured on a corpus the vendor chose: particular models, particular lengths, particular kinds of human writing, usually unedited. Your document was not drawn from that corpus. The further it sits from the test set — edited, mixed, translated, unusual — the less the number says about it.
Averages hide the tail
A high average accuracy is compatible with concentrated failure: short texts, non-native English, human writing that happens to be very uniform. The cases where detection goes wrong are not spread evenly — they cluster on particular writers and particular documents, and an average has no way to warn you whether yours is one of them.
The target moves
Every new model release and every new humanizer shifts the distribution detectors are trained against. A figure measured before a model existed cannot cover it. Whatever number was true on the day of measurement is a snapshot, and the field it describes changes monthly.
The practical consequence: a decimal on a landing page cannot help you with the document in front of you, but evidence can. A report that shows which sentences read generated, which read human, and how strongly, gives you something to check against your own knowledge of the text — and something concrete to discuss when the verdict is challenged. That is the trade this page is offering: the confidence lives in the report, not the marketing.
Three verdicts
Named calls,
sentence-level receipts
Every scan commits to one of three verdicts, and every sentence carries its own score on a four-step scale, shaded in place in your text. The mixed document — the one binary tools round away — has its own name.
Human-Written
The rhythm, vocabulary, and structure sit inside the human range — a named call the engine will stand behind, with the sentence scores visible so you can see why.
AI-Assisted
Both signatures present: model and person, in either order. A binary human-or-AI verdict has nowhere honest to put this case — and on real documents it is the most common truthful answer.
AI-Generated
Model output start to finish, including text already run through a humanizer — the evasion route the engine is specifically trained against.
Evaluate for yourself
The ten-document test
beats any landing page
You do not have to take any vendor's word — ours included. The evaluation that actually predicts how a detector will treat your text takes about fifteen minutes and requires nothing but documents you already have.
Gather known answers
Collect writing whose origin you are certain of: several pieces of your own from before 2023, something a colleague wrote, and a few texts you generate on purpose — one raw, one you edit heavily, one pushed through a paraphraser.
Run them cold
Paste each into the detectors you are comparing. Record the verdicts before checking against what you know — reading the answer first makes every result feel confirmed.
Score the directions
Count the two failure types separately: AI text called human, and your human text called AI. Which failure costs you more depends on what you are protecting — and no single advertised accuracy figure can price that trade for you.
What we will not claim
Why there is no number here
The absence of an accuracy figure on this page is the point of this page, so to be explicit: we publish no head-to-head benchmark against Winston AI, and no standalone accuracy percentage for markhuman, because we cannot produce one that would honestly transfer to your document. No detector is perfect. Short passages, non-native English, and heavily edited drafts are hard for the entire field, ours included. A report — anyone's report — is evidence to weigh, not proof.
The full account of what the report can and cannot tell you is on the AI detector page. The Originality.ai and ZeroGPT pages make the same case against other corners of the market.
FAQ
Questions about switching
Accuracy claims, how to run your own evaluation, and what to do when detectors disagree.
Is markhuman a free Winston AI alternative?
The scanner on this page is free without an account or card: 50 scans a day, up to 2,500 words per scan, 50-word minimum before the engine will score. Paid plans exist for longer documents and monthly volume; nothing in the report itself is gated.
Winston AI advertises 99.98% accuracy. Why doesn't markhuman quote a number?
Because the honest version of that number needs a paragraph of caveats: measured on which test set, which models, which text lengths, at which threshold, on which date? Winston's figure is its own measurement on its own terms — as is every vendor's headline number, and as any number of ours would be. We would rather show you the per-sentence evidence for each verdict than ask you to trust a decimal. When you evaluate any detector, including this one, ask how the number was produced; if the answer is not public and specific, treat it as marketing.
So how should I judge whether a detector is any good?
Test it on text you already know the truth about — your own writing, a colleague's, and something you generated on purpose. Ten known-answer documents tell you more about how a tool behaves on your kind of text than any advertised figure. Pay attention to the failure directions separately: a tool that never flags AI text and a tool that flags your human writing are both failing, but in ways that matter differently depending on what you are protecting.
What does the three-verdict system change in practice?
It gives the mixed document a name. Most real text now involves a model somewhere — a draft, an edit, a rewrite — and a binary human-or-AI call has to round that case toward one pole or the other. AI-Assisted is a positive finding with its own evidence, which means a policy, an editor, or a teacher can respond to the actual situation instead of a rounded one.
Will markhuman and Winston AI agree on the same text?
Not always — no two detectors reliably agree, because they are different classifiers trained on different corpora with different thresholds. Disagreement between tools is a property of the field. When it happens, the passage is genuinely ambiguous; the sentence-level shading here shows which specific lines carry the signal, which is the most useful thing a detector can hand you at that point.
Is my text stored, or used to train the model?
No. Text is scored in memory and never stored. Given what people paste into detectors — unpublished manuscripts, student work, client documents — retention policy is worth checking on every tool you use, before the text is in the box.
Run the ten-document test on us
Text you already know the truth about is the only benchmark that transfers. Free — 2,500 words per scan, 50 scans a day, no card.