AI IN SCHOOLS
AI detectors falsely flagged 61.2% of essays by non-native English writers. AI detectors falsely flagged 61.2% of essays by non-native English writers.
Why detectors fail this way
They do not recognise machine authorship. They largely measure how statistically predictable a piece of writing is — and a writer with a smaller working vocabulary produces more predictable sentences. That is the same signal the tool reads as "machine".
The study confirmed the mechanism directly: adding more literary language to non-native writing reduced misclassification, and simplifying native writing increased it. The tools are penalising limited linguistic range, not detecting AI.Same study. Independent benchmark work (RAID, ACL 2024, 6.2 million generations across eleven models and eleven adversarial attacks) separately found detectors generalise poorly to unseen models and lose accuracy from changes as simple as swapping words for synonyms.
A tool that misclassifies the majority of one group's genuine work cannot support an accusation against an individual. It may be usable as a soft signal across a whole cohort. It is not evidence about a person — and using it as evidence falls hardest on students already at a disadvantage: international students, migrants, and anyone still building fluency in the language they are assessed in.
For teachers
Do not open with the score
An accusation that begins "the detector says" puts a student in the position of proving a negative against a tool whose error rate they cannot examine. Whatever the outcome, the relationship is damaged.
Assess the process, not just the product
Drafts and version history. Some writing done in class. A short oral conversation about the submitted work. Assignments tied to a specific class discussion, a local example, or this week's material — things a general-purpose model has no access to.
These are not merely anti-AI measures. They measure understanding better than a finished essay does, and they were good practice before any of this existed.
Set the expectation, not the trap
State plainly what is allowed: research, grammar checking, structuring — and what is not: submitting generated prose as your own. Ask students to declare what they used. Most will, if declaring is not punished by default.
For students
Ask what the evidence is
If it is a detector score, ask a specific question: what is this tool's false-positive rate for writers with my language background? The published research says the honest answer is often uncomfortable.
Then supply process evidence: drafts, document version history, notes, search history, timestamps. And ask to talk about the work — being able to explain your argument, defend a choice and extend it live is far stronger evidence of authorship than any score.
Say so, early and specifically
"I used it to check grammar and to help structure section two" is a survivable conversation. Being caught after denying it is not. Institutions are far more forgiving of disclosed use than of concealment, because concealment is the part that looks like dishonesty.
For parents
Two things are worth knowing before you take a side. Detector scores are not proof — you are entitled to ask what the accusation rests on and what that tool's error rate is for your child specifically. And the skill worth protecting is not essay production; it is the ability to think through a problem and defend a conclusion. A student who can explain their own work is fine. A student who cannot is in trouble whether or not a machine wrote it.
THE WIDER GAP
UNESCO's 2025 global mapping found that 171 of 194 member states reference media and information literacy in national policy, but only 17 countries — fewer than 9% — have a stand-alone policy for it. Recognition is nearly universal; implementation is rare. Schools are being asked to handle this with tools that do not work and training most of them have not been given.UNESCO, Media and Information Literacy for All: Closing the Gaps, 2025. Note: 17 is a count of countries, not a percentage — a distinction this site got wrong once and corrected in public.
The honest position is uncomfortable for both camps. Students using AI to skip the thinking are cheating themselves out of the one thing school is for. And schools using detectors to accuse individuals are doing something worse — punishing people, disproportionately the ones already disadvantaged, on evidence that peer review says is unreliable. The way through is not better detection. It is assessment that makes concealment pointless: ask them to think out loud.