Loading...
Loading...
A student wrote a TOEFL practice essay about renewable energy. Simple, structured, clear. Four paragraphs. Five hundred words. GPTZero flagged it as 89% AI-generated. The student had never used ChatGPT. They had written it in a coffee shop with a notebook and a pen before typing it into Google Docs.
When people hear this story, the usual reaction is that the detector is broken. It is not. The detector is doing exactly what it was designed to do. The problem is that what it was designed to do, and what people think it does, are two different things.
This guide explains the two statistical signals almost every commercial AI detector runs on, why those signals systematically misfire on TOEFL-style writing, and what that means if you write in English as a second language. It is the technical companion to our complete guide to AI detector bias against non-native English writers.
Despite the marketing language, AI detectors do not detect AI. They detect statistical patterns that happen to correlate with machine-generated text. The two patterns that dominate the category are perplexity and burstiness. Almost every major detector, including GPTZero, Originality.ai, Turnitin, and Copyleaks, builds its core classifier on these two signals, sometimes augmented with proprietary neural classifiers but rarely replacing them.
Perplexity is a score that language models use to measure how surprising a piece of text is. Imagine reading a sentence one word at a time. At each step, a language model predicts the most likely next word. If the actual next word is the one the model would have guessed, that word has low perplexity. If the actual next word is unusual or unexpected, it has high perplexity.
The mathematical intuition is straightforward. If a writer always picks the most obvious next word, their text scores low on perplexity. If a writer picks unusual words, their text scores high. Detectors treat low perplexity as a signal of AI generation because LLMs are literally trained to pick the most probable next token. Statistically, LLM output clusters at the low end of the perplexity distribution.
Burstiness measures variation in perplexity across a document. Human writing tends to burst. A long, complex sentence with several clauses is followed by a short, punchy one. A surprising word choice sits next to a common one. The variation is irregular and human. LLM writing tends to be more uniform. Sentences have similar length and similar structure, creating a flatter statistical profile.
Detectors treat low burstiness as a second signal of machine generation. Combined, low perplexity and low burstiness are the two flags that drive most AI scores above the threshold.
Here is where the design of detectors collides with the design of TOEFL essays.
TOEFL is a standardized test of academic English. Test takers are explicitly trained to write in a specific way: clear thesis, structured paragraphs, conventional transitions, predictable vocabulary. The whole point of TOEFL writing instruction is to produce text that is unambiguous, grammatical, and easy to follow. Every choice a TOEFL writer makes, simpler vocabulary, standard syntax, conventional paragraph structure, pushes their writing toward low perplexity and low burstiness.
That is the same statistical signature detectors flag as AI-generated. The detector is not lying about what it sees. It sees text that statistically resembles machine output. It is wrong about the cause. The text resembles machine output because clear, careful ESL writing and clear, careful machine writing both converge on the same statistical profile. They get there by different paths, but the math cannot tell the paths apart.
This is the structural reason the bias persists across years and across detectors. It is not a bug in one company's classifier. It is a consequence of using perplexity and burstiness as proxies for human authorship when those metrics are also proxies for linguistic caution.
The 2023 Stanford study by Liang et al., published in *Patterns*, made the mechanism undeniable. The team ran 91 TOEFL essays through seven detectors. The average false positive rate was 53.5%. One detector flagged 97.8% of the essays as AI-generated. At least one detector flagged every single TOEFL essay.
The decisive experiment was a follow-up. The team used ChatGPT to "enhance" the vocabulary of the TOEFL essays, swapping simpler word choices for more native-sounding phrasing. The authorship did not change. Only the surface vocabulary changed. The average false positive rate dropped by 49.7 percentage points, from 53.5% to 3.8%.
Read that result carefully. The same essays, written by the same people, were flagged as AI when they sounded like non-native English and not flagged when they sounded more native. The detectors were not detecting AI. They were detecting the absence of native-level fluency.
Many ESL writers use Grammarly, QuillBot, or ChatGPT to polish grammar before submission. This is a reasonable practice. It also pushes false positive rates higher.
Polishing tools smooth out irregularities. They replace unusual phrasings with conventional ones. They normalize sentence length. Each of those changes lowers perplexity and lowers burstiness. After a polish pass, your writing is statistically closer to machine output than it was before, even though you wrote every word.
The research on this is consistent. In 2026 tests, ESL essays polished with QuillBot showed false positive rates 12 to 18 percentage points higher than the same essays before polishing. The polish tools are doing exactly what detectors reward, which is the opposite of what writers using them expect.
If you are a non-native English writer, three practical implications follow from this mechanism.
First, your false positive risk is not random. It is structural. You cannot eliminate it by being a better writer. You can reduce it by being a more varied writer, but the baseline risk is higher for you than for a native speaker writing the same kind of essay.
Second, the detector's score is not evidence about you. It is evidence about the statistical profile of your text. The two are not the same thing, and the gap between them is the entire basis for an appeal.
Third, the choice to use polish tools is a real trade-off. Cleaner grammar versus higher detection risk. There is no free lunch. If you polish, keep your unpolished draft. If you are flagged, the unpolished version is evidence that the polish, not you, is what moved the statistical profile.
You can shift your perplexity and burstiness profiles without changing what you actually have to say. The techniques are not tricks. They are habits that good writers in any language use.
Vary your sentence length. If you have written three sentences of roughly the same shape in a row, break one up or combine two. The variation is what raises burstiness.
Use specific vocabulary where you can. Generic abstract nouns, "things," "people," "ideas," lower perplexity. Specific nouns, "the 2023 IPCC synthesis report," "the Stanford Liang study," raise it.
Allow some informal rhythm. A short sentence. A fragment. A question. These raise burstiness without affecting meaning.
Read your draft aloud. If every sentence sounds the same length and weight, you have a burstiness problem. If you stumble, the reader will too, and so will the detector.
For a fuller playbook, see our ESL writing tips to reduce AI detection false positives.
None of these techniques fully eliminate the bias. A non-native English writer who varies sentence structure and uses specific vocabulary will still be flagged more often than a native speaker doing the same. The structural bias is in the metric, not in the user.
This is why the deeper fix has to come from institutions, not from individual writers. Detectors that flag clear ESL writing as AI are not fit for purpose in academic integrity decisions. Until they change, your job is to understand the mechanism, write in ways that reduce your statistical exposure, and know how to defend yourself when the system gets it wrong.
*This is the technical companion to our complete guide to AI detector bias against non-native English writers. If you have already been flagged, read How to Appeal an AI Detector False Positive: A Step-by-Step Guide for ESL Students. For a broader primer on detection mechanisms, see our guide to how AI content detectors work.*
Humanize AI text to sound naturally human with EvalHub.
Start Free Trial