Why AI Detectors Flag Real Student Writing

A detector does not measure whether you used AI. It measures how predictable your writing is, and ordinary human writing is often very predictable. Here is why real students get flagged, and how to protect yourself.

Jackson Wright7 min read
Two different writing pages casting the same shadow onto an AI-detection warning sensor

Listen while you read

0% read7 min left

Listen to this article
0:00--:--
On this page

In a study published in the Cell Press journal Patterns, seven widely used AI detectors were shown ninety-one essays written by real people, all of them non-native English speakers. The detectors labelled more than half of that real human writing as "AI-generated". If you have been flagged and you know you wrote your work, start with this fact: a detector does not measure whether you used AI. It measures how statistically predictable your writing is, and plenty of careful, ordinary human writing is very predictable.

What a detector actually measures

AI detectors do not read for meaning. They score two surface patterns. The first is perplexity: how surprised a language model is by your next word. Writing that is plain, formulaic, or built from phrases a model has seen often has low perplexity, and detectors read low perplexity as machine-written. The second is burstiness: how much your sentence length and complexity vary. People tend to vary; models tend to be steady, so steady human writing looks suspicious too.

Read those two together and the problem is obvious. The two things that get real students flagged are writing simply and writing consistently. Those are the exact habits that exam preparation and clear-writing advice train into you. The better you are at writing a plain, well-structured, uniform response, the more a detector may mistake you for a machine.

Who gets flagged the most

The Patterns study found the pattern falls hardest on people with a narrower range of English expression.

61%of real non-native essays flagged as AI, on average
19.8%flagged by all seven detectors at once
97.8%flagged by at least one detector

The proof that the detectors were keying on style, not authorship, is what happened next. When the same human essays were rewritten with richer, more varied language, the false-positive rate fell from about 61% to under 12%. Nothing about who wrote them changed. Only the writing style did. The authors of the study urged that detectors not be used in evaluative or educational settings, especially for non-native English speakers.

Even the makers cannot do it reliably

The company best placed to detect AI writing is the one that makes the AI. In January 2023 OpenAI released its own classifier to spot AI-written text. It withdrew the tool that July, citing a low rate of accuracy. On OpenAI's own figures the classifier correctly identified only about a quarter of AI-written text, while falsely flagging human writing roughly nine times in a hundred, and it was unreliable on anything under about a thousand characters.

The best-known school detector, Turnitin, scopes its own confidence tightly. Its widely quoted false-positive rate of under one percent applies only to documents it judges more than one-fifth AI-written; for documents below that threshold it now shows no score at all, because false positives there are higher. Its own chief product officer has put the sentence-level false-positive rate at around four percent. Even taken at face value, a one-in-a-hundred document rate across a whole year group means real students get flagged.

What this means for you

A flag is not a finding. It is a number from an inconsistent tool, and the people who run the system in Queensland say so. The QCAA's own guidance states that AI detection tools are inconsistent and should be used with caution and discernment, and its academic-integrity approach asks schools to authenticate a student's own work through a range of strategies, such as drafts, checkpoints, and version history, rather than leaning on a detector's score.

That is the important shift. The QCAA framework is built on demonstrating authorship, not on trusting software. Every QCE student completes an academic-integrity course and must submit original work, accurately acknowledging any contribution from other sources, including AI. Using AI to generate work you submit as your own is a breach. Being flagged by a detector, when you wrote the work yourself, is not.

If you are accused

The defence is evidence that you did the work, and it is the same trail the QCAA authentication approach expects you to have anyway. Build it as you write, not after.

  • Leave version history switched on. Google Docs and Microsoft Word record every edit with a timestamp, which shows the work forming over time.
  • Keep your drafts, planning notes, outlines, and annotated research rather than deleting them.
  • Keep the trail around the writing: your search history, library loans, and any teacher checkpoints.
  • If you are questioned, stay calm and produce the trail. Ask what evidence beyond a detector score is being relied on.
  • Do not try to beat the detector by inflating your language. Write your own work and keep the proof that you did.

Key takeaways

Detectors measure predictability rather than authorship, they are unreliable by their makers' own figures, and a process trail is the real protection.

PointDetails
A flag is not proofDetectors score perplexity and burstiness, which measure writing style, not whether AI was used.
Plain writing gets flaggedSimple, consistent, exam-trained writing looks machine-like, and non-native English writers are hit hardest.
Even the makers gave upOpenAI withdrew its own detector for low accuracy; vendors scope their low error rates tightly.
QCAA relies on authenticationThe QCAA calls detectors inconsistent and asks schools to verify authorship through drafts and process.
Keep the trailVersion history, drafts, and notes are the evidence that you wrote your own work.

What I tell students who are scared

The messages that worry me most are from students who did nothing wrong and have been told a machine says otherwise. The fear is real, and it is not irrational, because the accusation feels total and the tool sounds authoritative.

What steadies it is understanding what the tool is. It is a style gauge with a known habit of misfiring on plain, ordinary writing. It is not a lie detector. When you see it that way, the flag stops being a verdict and becomes a claim you can answer, with drafts, with edit history, with the ordinary evidence of having done the work.

Write in your own voice, keep your working, and do not let a probability score convince you that your own work was not your own. A flag does not mean you did anything wrong.

Feedback that improves your work, not detection

ISMGenius is not a detector, and it does not write your assignment. It reads your draft against your task's own ISMG and gives you rubric-aligned feedback, criterion by criterion, so you can improve your own work before you submit it. It is built for Queensland senior students who want an informed estimate of where a draft sits, not a shortcut around doing the work.

A gauge labelled predictability rather than authorship, showing plain human writing landing in the same zone as AI-generated text
Detectors score how predictable writing is. Plain human writing can land in the same zone as machine text.

You can read how ISMGenius approaches academic integrity and AI tools, see how originality scanning is explained in the help centre, and run a draft for rubric-aligned feedback. The QCAA sets out its position on artificial intelligence in schools and the academic integrity requirement, and the false-positive findings come from a study on how detectors are biased against non-native English writers.

Frequently asked questions

Why did an AI detector flag my essay when I wrote it myself?

Detectors score how statistically predictable your writing is, using measures called perplexity and burstiness. Plain, consistent, exam-trained writing looks predictable, and detectors read that as machine-written. The flag reflects your writing style, not whether you actually used AI.

How often do AI detectors get it wrong?

A study in the journal Patterns found seven detectors flagged more than half of real essays by non-native English writers as AI-generated, on average about 61%. OpenAI withdrew its own detector for low accuracy, and it falsely flagged human writing about nine times in a hundred.

Does a flag from an AI detector prove I cheated?

No. A flag is a number from an inconsistent tool, not a finding. The QCAA itself describes AI detection tools as inconsistent and says they should be used with caution. Its academic-integrity approach relies on authenticating your work through drafts and process, not on a detector score.

What should I do if I am accused of using AI?

Stay calm and produce your process trail: version history, drafts, planning notes, and research. Ask what evidence beyond the detector score is being relied on. Keeping version history switched on and planning in the same document as your essay is the strongest proof of authorship.

Who is most likely to be wrongly flagged by an AI detector?

Non-native English writers are hit hardest, because their writing tends to use a narrower range of expression, which detectors read as machine-like. In one study, rewriting the same human essays with richer language dropped the false-positive rate from about 61% to under 12%.

Back to all articles

Read next

Five subject result blocks passing through a curved scaling lens and stacking into a single ranked column

How the Queensland ATAR Is Calculated: Scaling and the TEA

Your ATAR is not your subject marks added up. Two organisations, three stages and one scaling step decide it, and the scaling step is the one almost everyone misreads. Here is the whole process, with QTAC's own published numbers.

Jackson Wright6 August 2026
A Queensland Certificate of Education built from five stacked blocks, with a newly added fifth block for academic integrity

The QCE Academic Integrity Requirement, Explained

A fifth condition now sits between you and a QCE, and it takes about an hour. Here is what the academic integrity requirement asks for, who it applies to, and the eleven kinds of misconduct behind it.

Jackson Wright6 August 2026
A dense alphabetical grid of uniform word blocks with one block lifted forward and magnified in coral

All 117 QCAA Cognitive Verbs: The Complete List

QCAA's cognitive verb framework runs to 117 terms, and your task sheet draws its first word from it. Here is the complete list, searchable, with the official wording for every entry.

Jackson Wright6 August 2026