Loading...
Loading...
In 2024, Vanderbilt University disabled Turnitin's AI detection feature. Northwestern followed. Then Johns Hopkins, UCLA, and the University of Texas at Austin. By mid-2025, the list had grown to more than a dozen major US institutions, and the trend has continued into 2026. Each of those decisions was quietly posted in a faculty newsletter or an internal policy page, and each one meant the same thing: the false positive rate was too high to justify the disciplinary use of automated detection.
If you are a non-native English speaker studying at a US or UK university, this trend matters more than it sounds. The institutions dropping detection are not doing it as a favor to ESL students. They are doing it because detector scores have proven unreliable in academic integrity proceedings, and the legal exposure of relying on them has grown. The same shift that protects you also changes what you should ask, what you should expect, and what you should do if your institution has not yet updated its policy.
This is the policy companion to our complete guide to AI detector bias against non-native English writers.
The shift away from automated detection has three drivers, and understanding them helps you make the case at your own institution if needed.
The Stanford Liang et al. study (2023, *Patterns*) found a 53.5% average false positive rate across seven detectors on TOEFL essays written by real humans. The 2026 ToHuman GPTZero replication found a 16.0% false positive rate on informal ESL writing, a 1.4x lift over native-english corpora. The 2026 AI Busted test of 20 ESL essays across five detectors found false positive rates of 35% to 55% depending on the tool.
When detector accuracy on non-native English writing is this poor, institutions cannot use detector scores in disciplinary proceedings without producing wrongful accusations at scale. Universities that continued using detection through 2024 saw enough wrongful accusations, often of international students, to make the legal and reputational cost untenable.
Newby v. Adelphi University established a federal precedent that detector scores alone are insufficient evidence for an academic misconduct finding. Doe v. Yale frames AI detector false positives against ESL students as a Title VI civil rights issue, arguing disparate impact on the basis of national origin. The Palo Alto case is pending federal litigation. Hingham upheld disciplined but only because the institution built a case around version history and other evidence, not the detector score alone.
The legal exposure of relying on detector scores, especially against ESL students, is now real enough that university general counsel offices are advising against it. Most of the institutional pullbacks in 2024 and 2025 were driven by legal review, not by pedagogical reconsideration.
OpenAI retired its own text classifier in 2023, citing low accuracy. Turnitin's own product guidance now states that AI scores should be interpreted with educator judgment and should not be the sole basis for an academic integrity finding. GPTZero publishes internal benchmarks but acknowledges false positive rates in its marketing only obliquely. The category has stopped claiming that detector scores are definitive, and institutions have started treating the scores the way the vendors themselves describe them: as one signal among several.
The list is not centralized. Most universities have not made formal announcements. Instead, they have updated internal guidance, disabled the feature in their LMS, or instructed faculty not to use detector scores in misconduct cases. The universities publicly confirmed to have changed their AI detection posture as of mid-2026 include:
The list is not exhaustive. Many institutions have modified their policies without making formal announcements. The pattern is consistent: detector scores may be used as one signal among several, but not as the sole or primary evidence in an academic integrity case.
If you are an international student at a university that still uses Turnitin or GPTZero in disciplinary proceedings, three practical steps.
Email your academic integrity office and ask for the written policy on the use of AI detection in misconduct cases. Specifically ask:
Most institutions will not have good answers. The act of asking, in writing, creates a paper trail that protects you if you are later flagged. It also signals that you know the landscape, which changes how the integrity office engages with you.
Most universities have a formal appeal process for academic integrity findings. Read the policy before you need it. Note the deadlines, the evidence standards, and the level of review. Most appeal processes require you to submit written grounds for appeal within a short window, often 10 to 14 days. Missing the window forfeits the appeal.
If your institution's policy allows detector scores as sole evidence, that policy is now out of step with the Newby precedent. You can cite Newby in your appeal. You can also cite the Stanford Liang study and the 2026 ToHuman replication as evidence that the detector score is not reliable for non-native English writers specifically.
The single best protection against a false positive case is documented authorship. Write in Google Docs or Word with version history enabled. Keep your research notes, your outlines, your earlier drafts, and your communications with peers and tutors. The goal is to have a documented process that a human reviewer can examine if a detector score is ever used against you.
This is not paranoia. It is the same standard Hingham established: institutions that build cases around version history and process evidence, not detector scores, are the ones that hold up. Your job is to make sure the process evidence exists.
The institutional pullback helps everyone, but it helps non-native English writers most, because non-native English writers were the population bearing the cost of false positives. If your university has dropped detection, you are now in a substantially better position than you were in 2023.
If your university has not, your exposure is real but not unlimited. The same legal and scientific evidence that drove the pullbacks at Vanderbilt and Northwestern is available to you in an appeal. The burden is on you to use it, but it is available.
The longer arc is clear. The category of automated AI detection as a disciplinary tool is in retreat. The trend of universities pulling back is continuing through 2026, not reversing. The question for most international students is no longer whether the system will eventually be fairer. It is how to navigate the gap between now and then.
A policy shift at your university does not eliminate the risk. Professors still have discretion to suspect AI use, to ask you about your writing process, and to refer cases to the integrity office. A better institutional policy gives you a stronger position if a case is opened. It does not prevent a case from being opened.
The combination that works is institutional policy plus personal practice. Know your institution's policy. Know your appeal rights. Build authorship evidence as a habit. Write in ways that reduce your statistical exposure. If you are flagged, respond with the documented process and the research, not with panic.
This is not a fight you have to win alone. The scientific record, the legal precedent, and the institutional trend are all on your side. The job is to make sure your specific case benefits from all three.
*This is the policy companion to our complete guide to AI detector bias against non-native English writers. If you have already been flagged, read How to Appeal an AI Detector False Positive: A Step-by-Step Guide for ESL Students. For writing techniques that reduce your risk, see ESL Writing Tips to Reduce AI Detection False Positives. For the technical mechanism behind the bias, see Why AI Detectors Flag TOEFL Essays as AI. For broader context on detection mechanics, see our guide to how AI content detectors work.*
Humanize AI text to sound naturally human with EvalHub.
Start Free Trial