AI Detection Tools Don't Work. Here's What to Do Instead.
Detectors produce false accusations against real writing, especially from non-native English speakers. If you need to verify authorship, verify process instead.
By HIMEXA Editorial
Teachers, editors and hiring managers all want the same thing: a reliable way to tell whether a person wrote what they submitted. Detection tools promise exactly that, which is why they spread so fast.
They do not deliver it, and the way they fail causes real harm.
How detectors actually work
They measure statistical properties — how predictable each word is given the previous ones. Text that is unusually smooth and predictable scores as machine-written.
The problem is that plenty of human writing is smooth and predictable: formal academic prose, technical documentation, and writing by people composing carefully in a second language.
Who gets falsely accused
Disproportionately, non-native English speakers. Writing in a learned language often means simpler sentence structures and more conventional word choices — the exact profile detectors flag.
A tool that is wrong more often for one group of people is not a neutral tool. In an academic or professional setting, that is a serious problem, not a rounding error.
Why the arms race can't be won
Anyone deliberately evading detection can do so trivially — rephrase, add variation, run it through another model. Detection catches the careless, not the deceptive.
So the false positives fall on honest people while the intended targets pass. That is the opposite of what the tool is for.
Verify process, not text
For education: assess drafts, outlines and revisions rather than only the final artefact. Ask students to explain their choices verbally. Design assignments around personal experience, local context or in-class material that a general model cannot supply.
For hiring: use a live exercise and a conversation about the work. Someone who understands their own submission can discuss it; someone who cannot, cannot.
For editing: ask for sources and reasoning. Verification of claims is more useful than provenance of sentences, and it catches the failure mode that actually matters — content that is wrong.
If you must use a detector
Treat a flag as a prompt to have a conversation, never as evidence. Never confront someone with a score as proof. And never apply it in a way that puts more burden on people writing in a second language.
The honest position is that we do not currently have a reliable authorship detector, and pretending otherwise damages people who did nothing wrong.