The most cited claim in Turnitin AI detection cases is that the tool has a 1% false positive rate. That figure comes from Turnitin's own testing. Independent research, including work by University of Maryland computer scientists Soheil Feizi and colleagues, has found substantially higher error rates and a set of structural weaknesses that any student facing a Turnitin accusation should understand.
What the Maryland research actually tested
The most widely discussed work from the University of Maryland on AI detection is the 2023 paper Can AI-Generated Text be Reliably Detected? by Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. The paper appeared as a preprint in March 2023 and was subsequently revised. It tested how AI text detectors perform against three specific attacks: paraphrasing, recursive paraphrasing, and adversarial edits to AI-generated output.
The authors reached two conclusions that matter for accused students. First, light paraphrasing of AI output reliably defeats detection. Second, detectors face a mathematical trade-off: as the underlying language models improve, the statistical gap between AI and human writing shrinks, which forces detectors to either raise false positive rates or miss more AI text. There is no free lunch. The paper's authors concluded that reliable detection may be fundamentally hard as models scale.
Follow-up work by researchers at Maryland and other institutions continued to probe these limits through 2024. For a broader look at the peer-reviewed record, see our summary of 2023 research on AI detector accuracy.
What Turnitin itself claims, and where those claims break down
Turnitin has publicly stated a document-level false positive rate below 1% and a sentence-level false positive rate around 4%. Turnitin's own guidance, published on its site, has also cautioned that the AI writing indicator is not designed to be the sole evidence in an academic misconduct case and that scores below 20% AI content carry an elevated false positive risk. We cover the company's own acknowledgments in detail in what Turnitin has publicly said about false positives.
The gap between Turnitin's vendor testing and independent findings has two main causes. Turnitin's internal validation uses text drawn from a fixed distribution: essays that resemble the training data. Real student writing spans a wider range: non-native English, formal academic register, technical prose, edited output from grammar tools. Studies by Weber-Wulff et al. (2023) in the International Journal of Educational Integrity and Liang et al. (2023) in Patterns both found error rates well above vendor claims on real-world writing samples.
Why a small false positive rate still produces many wrongly accused students
Even if Turnitin's own 1% document-level false positive rate were exactly correct, the volume of submissions makes wrongful flags common. Turnitin has reported processing tens of millions of AI-scanned submissions. A 1% rate applied to that base produces hundreds of thousands of false flags. A 4% sentence-level rate, which Turnitin also acknowledges, means most long documents will contain at least some incorrectly flagged sentences.
This is why the research consensus is not that detectors are useless, but that they cannot support a misconduct finding on their own. A probabilistic score with a non-trivial error rate is a screening signal, not evidence.
How to use the Maryland research in your response
If your case rests on a Turnitin AI score, the Maryland paper supports three specific arguments you can make in a written response or hearing:
- Detection is not deterministic. Peer-reviewed and preprint research by Sadasivan, Feizi, and colleagues at the University of Maryland shows AI detection faces a fundamental accuracy ceiling as language models improve.
- Vendor false positive rates understate real-world error. Independent testing consistently finds higher error rates than Turnitin's internal validation on writing that falls outside typical training distributions.
- A score is not evidence of a policy violation. Turnitin itself has told institutions the AI indicator should not be the sole basis for a misconduct finding.
Pair those research points with concrete process evidence: Google Docs version history, browser history, notes, and prior drafts. The procedural rights FAQ covers what you are entitled to request from the institution before a hearing, including the exact detector version and score. If you are preparing a written response, NotBot generates a personalized defense package that names the detector, cites the applicable research including the Maryland work, and documents your writing process, ready in about a minute.
What this research does not prove
The Maryland work does not prove that your specific paper was misclassified. It shows the tool is unreliable enough that a score alone should not be treated as evidence. That distinction matters in a hearing. Your job is not to argue the detector is always wrong, but to establish that its output, combined with the specific facts of your case, does not meet the evidentiary standard your institution's policy actually requires.
Build your defense package
A personalized response that cites the Maryland research and documents your writing process, ready in minutes.
Get your defense package$49 one-time · Generated in 60 seconds