In a study published in the Cell Press journal Patterns, seven widely used AI detectors were shown ninety-one essays written by real people, all of them non-native English speakers. The detectors labelled more than half of that real human writing as "AI-generated". If you have been flagged and you know you wrote your work, start with this fact: a detector does not measure whether you used AI. It measures how statistically predictable your writing is, and plenty of careful, ordinary human writing is very predictable.
What a detector actually measures
AI detectors do not read for meaning. They score two surface patterns. The first is perplexity: how surprised a language model is by your next word. Writing that is plain, formulaic, or built from phrases a model has seen often has low perplexity, and detectors read low perplexity as machine-written. The second is burstiness: how much your sentence length and complexity vary. People tend to vary; models tend to be steady, so steady human writing looks suspicious too.
Read those two together and the problem is obvious. The two things that get real students flagged are writing simply and writing consistently. Those are the exact habits that exam preparation and clear-writing advice train into you. The better you are at writing a plain, well-structured, uniform response, the more a detector may mistake you for a machine.
Who gets flagged the most
The Patterns study found the pattern falls hardest on people with a narrower range of English expression.
The proof that the detectors were keying on style, not authorship, is what happened next. When the same human essays were rewritten with richer, more varied language, the false-positive rate fell from about 61% to under 12%. Nothing about who wrote them changed. Only the writing style did. The authors of the study urged that detectors not be used in evaluative or educational settings, especially for non-native English speakers.
Even the makers cannot do it reliably
The company best placed to detect AI writing is the one that makes the AI. In January 2023 OpenAI released its own classifier to spot AI-written text. It withdrew the tool that July, citing a low rate of accuracy. On OpenAI's own figures the classifier correctly identified only about a quarter of AI-written text, while falsely flagging human writing roughly nine times in a hundred, and it was unreliable on anything under about a thousand characters.
The best-known school detector, Turnitin, scopes its own confidence tightly. Its widely quoted false-positive rate of under one percent applies only to documents it judges more than one-fifth AI-written; for documents below that threshold it now shows no score at all, because false positives there are higher. Its own chief product officer has put the sentence-level false-positive rate at around four percent. Even taken at face value, a one-in-a-hundred document rate across a whole year group means real students get flagged.
What this means for you
A flag is not a finding. It is a number from an inconsistent tool, and the people who run the system in Queensland say so. The QCAA's own guidance states that AI detection tools are inconsistent and should be used with caution and discernment, and its academic-integrity approach asks schools to authenticate a student's own work through a range of strategies, such as drafts, checkpoints, and version history, rather than leaning on a detector's score.
That is the important shift. The QCAA framework is built on demonstrating authorship, not on trusting software. Every QCE student completes an academic-integrity course and must submit original work, accurately acknowledging any contribution from other sources, including AI. Using AI to generate work you submit as your own is a breach. Being flagged by a detector, when you wrote the work yourself, is not.
If you are accused
The defence is evidence that you did the work, and it is the same trail the QCAA authentication approach expects you to have anyway. Build it as you write, not after.
- Leave version history switched on. Google Docs and Microsoft Word record every edit with a timestamp, which shows the work forming over time.
- Keep your drafts, planning notes, outlines, and annotated research rather than deleting them.
- Keep the trail around the writing: your search history, library loans, and any teacher checkpoints.
- If you are questioned, stay calm and produce the trail. Ask what evidence beyond a detector score is being relied on.
- Do not try to beat the detector by inflating your language. Write your own work and keep the proof that you did.
Key takeaways
Detectors measure predictability rather than authorship, they are unreliable by their makers' own figures, and a process trail is the real protection.
| Point | Details |
|---|---|
| A flag is not proof | Detectors score perplexity and burstiness, which measure writing style, not whether AI was used. |
| Plain writing gets flagged | Simple, consistent, exam-trained writing looks machine-like, and non-native English writers are hit hardest. |
| Even the makers gave up | OpenAI withdrew its own detector for low accuracy; vendors scope their low error rates tightly. |
| QCAA relies on authentication | The QCAA calls detectors inconsistent and asks schools to verify authorship through drafts and process. |
| Keep the trail | Version history, drafts, and notes are the evidence that you wrote your own work. |
What I tell students who are scared
The messages that worry me most are from students who did nothing wrong and have been told a machine says otherwise. The fear is real, and it is not irrational, because the accusation feels total and the tool sounds authoritative.
What steadies it is understanding what the tool is. It is a style gauge with a known habit of misfiring on plain, ordinary writing. It is not a lie detector. When you see it that way, the flag stops being a verdict and becomes a claim you can answer, with drafts, with edit history, with the ordinary evidence of having done the work.
Write in your own voice, keep your working, and do not let a probability score convince you that your own work was not your own. A flag does not mean you did anything wrong.
Feedback that improves your work, not detection
ISMGenius is not a detector, and it does not write your assignment. It reads your draft against your task's own ISMG and gives you rubric-aligned feedback, criterion by criterion, so you can improve your own work before you submit it. It is built for Queensland senior students who want an informed estimate of where a draft sits, not a shortcut around doing the work.
You can read how ISMGenius approaches academic integrity and AI tools, see how originality scanning is explained in the help centre, and run a draft for rubric-aligned feedback. The QCAA sets out its position on artificial intelligence in schools and the academic integrity requirement, and the false-positive findings come from a study on how detectors are biased against non-native English writers.



