Ever since language models started writing entire essays, schools have been looking for a tool that settles the question at the push of a button: was this written by AI? Detector vendors promise exactly that, often backed by impressive percentages. The trouble is not only that those numbers rarely hold up in practice. It starts with the maths itself.
One percent sounds small until you scale it
Say a detector is right 99 percent of the time. Your school runs 1000 essays through it, and every single one was written honestly. A one percent false alarm rate then means ten students fall under suspicion, even though they did nothing wrong.

Even at 99 percent accuracy, ten out of 1000 honest essays come under suspicion.
The key point: false alarms grow with volume. The more texts you check, the more innocent students get flagged, even if the percentage stays the same. Each case is a difficult conversation with trust at stake. And 99 percent is a deliberately optimistic figure.
What the research shows
Real-world results are considerably worse. In 2023, a research team led by Weber-Wulff tested fourteen detection tools. Their verdict: the tools are neither accurate nor reliable.

Three research findings on AI detectors.
A second study, by Liang and colleagues, points to a fairness problem. English texts by people writing in a language that isn’t their first are flagged as AI especially often. The study looked at English essays, including TOEFL essays; it says nothing about other languages. The researchers suggest that detectors may unintentionally penalise writers with more constrained linguistic expression, and they explicitly caution against using these tools in evaluative or educational settings.
Even OpenAI, the company behind ChatGPT, shut down its own detector in July 2023, citing its low rate of accuracy.
What works better
A detector result is not proof. At most, it is a reason to look more closely. In everyday teaching, these three recommendations from the video go further:
- Make the writing process visible. Drafts, notes and intermediate versions become part of the submission. Students who developed a text themselves can show how they got there.
- Talk with the student about the text. Not as an interrogation, but as part of the assessment: Why this structure? What was hard?
- Design tasks that require their own thinking. For example, by connecting them to your own lessons, personal experience or a decision that has to be justified.

Three approaches that achieve more than a detector score.
An exercise for your staff meeting
My suggestion: estimate together how many written assignments are handed in at your school each semester. Then apply a one percent false alarm rate, the optimistic figure. The number you get is how many conversations you would have to have with honest students. After that, ask a better question: which tasks in our own teaching would make a detector unnecessary?
Trust doesn’t come from software. It comes from good tasks.
Frequently asked questions
How reliable are AI detectors?
Not very. In 2023, a research team led by Weber-Wulff tested fourteen detection tools and concluded that they are neither accurate nor reliable. OpenAI shut down its own detector in July 2023, citing its low rate of accuracy.
Is an AI detector result proof?
No, at most it is a reason to look more closely. Even at an optimistic 99 percent accuracy, ten out of 1000 honestly written essays would be flagged wrongly.
Are AI detectors biased against non-native writers?
For English, yes: according to a study by Liang and colleagues (2023), English texts by people writing in a language that isn't their first are flagged as AI especially often. The study did not look at other languages. The researchers caution against using these tools in evaluative or educational settings.
Sources
Reuse


