There is a special kind of modern absurdity in watching universities outsource academic judgement to software that claims to detect whether students outsourced their academic judgement to software.
We have now reached the point where a student can write most of an assignment themselves, perhaps use an AI tool for a small section, a bit of structure, or a paragraph they were stuck on, and then get flagged more aggressively than someone who dumped the entire prompt into an LLM and went to make coffee. That is not just ironic. That is weapons-grade irony. The kind you should store in a lead-lined container and only handle with tongs.
A recent paper, “How are combinations of human-written words and LLM-generated words by ChatGPT, Copilot, Gemini and Grammarly detected by Turnitin?”, takes a detailed look at how Turnitin behaves when fed different mixtures of human-written and AI-generated text. The study tested scripts ranging from fully human-written to fully LLM-generated, using ChatGPT, Copilot, Gemini, and Grammarly. It also looked at what happens when AI-generated text is pushed through “humanising” or paraphrasing tools such as QuillBot, EasyessayAI, and RyneAI.
The results are exactly the kind of thing that should make educators pause before treating AI detector scores as evidence. Not a hint. Not a warning light. Not a useful clue to be investigated. Evidence.
Because according to the paper, Turnitin is not simply “a bit inaccurate”. It is inaccurate in a way that creates a deeply uncomfortable incentive structure. When the text contains a small amount of AI-generated content, Turnitin may overestimate the AI score. When the text contains a lot of AI-generated content, Turnitin may underestimate it. In plain English: students who mostly write their own work may be made to look worse, while students who use AI heavily may be made to look less guilty than they are.
That is not an academic integrity system. That is a casino with a login page.
The detector that blinks at the wrong time
The paper tested 81 scripts with carefully constructed combinations of human-written and LLM-generated words. One script was 100% human-written. Others contained 5%, 10%, 15%, 20%, and so on, all the way up to 100% LLM-generated content.
The first interesting result is that Turnitin did not produce AI scores for the 100% human-written script. It also did not flag the scripts where only 5% or 10% of the text was AI-generated. In those cases, Turnitin showed its little “nothing to see here” face.
At first glance, that sounds reassuring. Great. It does not randomly accuse fully human text in this experiment, and it ignores very low levels of AI text. Wonderful. Put the champagne back in the fridge, though, because the story gets strange quickly.
Once the AI-generated content reached around 15% to 40%, Turnitin often detected it, but its score was usually higher than the actual amount of AI-generated text. In other words, a script with 15% AI-generated content could be reported as having more than 20% AI-generated content. A script with 20% AI could be reported as 25%, 29%, or 31%, depending on the model used.
That matters because many institutions do not treat these numbers as soft signals. They treat them as suspicious numbers. And in academia, “suspicious number” is often only one awkward meeting away from “please explain yourself in writing”.
Then comes the truly weird part. When the percentage of actual AI-generated words got higher, Turnitin’s scores often became lower than the real amount. At 100% LLM-generated text, the 4000-word scripts were not all detected as 100% AI-generated. The scores were 87% for Copilot, 60% for ChatGPT, 87% for Gemini, and 82% for Grammarly.
Let that sink in.
A fully AI-generated ChatGPT script was detected as 60% AI-generated.
Meanwhile, some partially human-written scripts with much smaller amounts of AI were reported as having more AI than they actually contained.
That means the tool can be harsher on mixed-authorship work than on fully machine-generated work. It can exaggerate the small offender and undercount the full robot buffet. It is like a tax system that audits people for buying one coffee with a company card while giving a polite nod to the guy loading an entire data centre into the back of a van.
The punishment problem
Now, to be fair, Turnitin itself does not punish anyone. It produces a score. The punishment comes when humans treat that score as something more than it is.
And that is the real issue.
The danger is not just that AI detection tools are imperfect. All tools are imperfect. Spam filters are imperfect. Antivirus engines are imperfect. My ability to estimate how much pasta to cook is absolutely imperfect and has been the cause of many unnecessary carbohydrate incidents.
The problem is that AI detector scores are often wrapped in an aura of technical authority. The interface looks official. The number looks precise. The institution has paid for it. The student is nervous. The teacher is busy. The process has a policy document. Suddenly, a probabillistic guess starts wearing a little judge’s wig.
This is especially dangerous because the number looks mathematical. “32% AI-generated” feels like a measurement. It has the same aesthetic as battery percentage or CPU load. But with AI text detection, we are not measuring a physical property. We are making an inference about authorship from statistical patterns in language. That is a much shakier thing.
Text does not carry a tiny “Made by ChatGPT” watermark stamped into every noun. AI detectors look for patterns. Predictability. Burstiness. Sentence structure. Token distributions. The kind of linguistic fingerprints that may indicate generated text, but can also appear in edited text, formulaic academic writing, non-native English writing, template-heavy assignments, or just the prose style of someone who has been trained to write in the beige concrete bunker we call “academic tone”.
Academic writing is already full of phrases that sound AI-generated because academic writing has spent decades training humans to write like underfunded language models.
“Furthermore, this paper seeks to explore…”
“In conclusion, the findings indicate…”
“It is important to note that…”
Half of academia already sounds like a committee prompted itself into existence.
Mostly human, suspiciously punished
The most troubling finding in the paper is not simply that Turnitin gets scores wrong. It is the direction of the wrongness.
The study found that when the percentage of actual LLM-generated words was lower, Turnitin’s AI score tended to be inaccurately higher. When the percentage of actual LLM-generated words was higher, Turnitin’s AI score tended to be inaccurately lower.
That is the worst possible shape for a disciplinary tool.
Imagine a speed camera that overestimates your speed when you drive 55 km/h in a 50 zone, but underestimates it when someone blasts past a school at 120. You would not call that “a useful contribution to road safety”. You would call it broken and demand to know who signed the procurement contract.
For students, the implications are ugly. A student who writes an essay themselves but uses AI to polish a weak section, generate an example, or rephrase something clumsy may end up with a score that looks worse than the actual usage. Another student who submits a mostly or fully AI-generated text may get a score that sounds less dramatic than the reality.
This is how you create perverse incentives.
If a system punishes messy, partial, honest, or clumsy AI use more visibly than polished, total AI use, then students learn the wrong lesson. They do not learn “write your own work”. They learn “if you are going to cheat, go all in and use better laundering”.
That is not integrity. That is training people in operational security.
The laundering layer
The study also tested what happens when fully AI-generated text is passed through so-called humanising tools. This is where things go from uncomfortable to farcical.
The paper reports that 100% Copilot-generated text humanised by QuillBot received a 0% AI score from Turnitin. In other cases, RyneAI made fully AI-generated text from Copilot, Grammarly, and Gemini appear human-written to Turnitin, also producing 0% AI scores. A ChatGPT-generated text humanised by RyneAI was detected as only 26% AI-generated.
This is not a minor edge case. This is the obvious next step for anyone trying to bypass detection. If an institution relies heavily on detectors, students who want to cheat will not simply paste raw ChatGPT output. They will generate, paraphrase, humanise, reorder, add a typo, sprinkle in a personal anecdote, and serve it with a garnish of fake hesitation.
The honest student, meanwhile, may use Grammarly, ChatGPT, or another tool in a limited way and then have to defend themselves against a number that cannot explain its own reasoning in a way that meets normal academic standards of evidence.
This is the part where universities should get very nervous.
Because when detection becomes the main enforcement mechanism, evasion becomes the main student skill. We have seen this before in every technological arms race. DRM creates cracking communities. Spam filters create better spam. Surveillance creates camouflage. AI detectors create AI text laundering.
Congratulations, we have invented academic money laundering, but for paragraphs.
“But Turnitin is one of the better ones”
Yes. That is precisely the problem.
Turnitin is widely used and often considered one of the more reliable AI detection systems. If one of the better-known tools behaves like this under controlled testing, then the correct response is not to shrug and say, “Well, it is probably good enough.”
Good enough for what?
Good enough to start a conversation with a student? Maybe.
Good enough to decide guilt? No.
Good enough to become part of a broader investigation that includes drafts, version history, oral defence, writing samples, assignment design, and human judgement? Possibly.
Good enough to automate suspicion at scale? Absolutely not.
This distinction matters. An AI detector can be a signal. It can be one input among many. It can say, “This text has statistical properties that are worth looking at.” That is a very different claim from “This student cheated.”
The first is a tool-assisted prompt for human review.
The second is a digital accusation.
The first belongs in a careful academic process.
The second belongs in the bin, preferably after being printed out and used to test whether the shredder also detects AI-generated confetti.
The deeper failure: pretending writing is still the same task
There is another uncomfortable truth hiding underneath all this: AI has changed what writing means.
For decades, written assignments were a proxy for learning. We asked students to write because writing demonstrated reading, understanding, structure, argumentation, and original thought. That worked reasonably well when producing a polished text required the student to personally perform most of those cognitive steps.
Now the relationship is more complicated.
A student can use AI to brainstorm, summarise, rephrase, outline, translate, critique, simplify, expand, and polish. Some of that is clearly useful learning support. Some of it is clearly cheating. A lot of it lives in the swampy middle where policy documents go to die.
Using AI to fix grammar is not the same as using AI to invent the argument. Using AI to ask “is this paragraph clear?” is not the same as asking it to write the paragraph. Using AI to generate three possible structures for an essay is not the same as submitting generated prose unchanged. But most detector systems do not know intent. They do not know process. They do not know whether the student learned anything.
They only see text.
And text is now the least reliable artefact in the whole process.
If we care about learning, we need to assess process again. Drafts. Notes. Version history. Reflections. Oral follow-up questions. In-class writing. Project logs. Reproducible work. Assignments tied to personal experience, local context, or recent classroom discussion. Not because these are impossible to fake, but because they move assessment away from the fantasy that a final polished document alone can tell us everything.
The detector tries to answer the question: “Was this text produced by AI?”
The better question is: “Can this student demonstrate the knowledge this assignment is supposed to measure?”
Those are not the same question.
Academic integrity cannot be outsourced
The temptation to outsource this problem is understandable. Teachers are overloaded. Institutions are under pressure. Students have access to tools that can generate passable essays in seconds. Nobody wants to manually investigate every suspicious submission.
But replacing human academic judgement with detector scores is not a solution. It is bureaucracy wearing a cybernetic hat.
The most dangerous version of this is when institutions quietly convert detector percentages into policy thresholds. Below this number, ignore. Above this number, escalate. Above that number, accuse. It feels efficient. It feels objective. It produces neat workflows. It also risks becoming deeply unfair.
A 25% score does not mean exactly one quarter of the assignment was written by AI. A 60% score does not mean the student wrote 40% themselves. A 0% score does not mean the student did not use AI. The study makes that painfully clear.
The score is not authorship. The score is a detector’s opinion about patterns in the text.
And detectors, like all opinionated machines, can be confidently wrong.
What universities should do instead
Universities should not ignore AI use. That would be naive. They should not pretend the old assignment model is still intact. That would be institutional cosplay. But they should also stop treating AI detection as a magic plagiarism microscope.
The sane approach is to separate academic integrity from detector worship.
First, institutions need clear AI-use policies that distinguish between assistance and substitution. “No AI” is easy to write and hard to enforce. “Use AI ethically” sounds nice but becomes meaningless unless students know what that means in specific assignments. A programming course, a literature essay, a medical reflection, and a management theory paper may all need different rules.
Second, students should be asked to declare AI use in a simple and non-punitive way when it is allowed. Not a confession booth. Not a legal affidavit. Just a normal part of academic transparency: which tools were used, for what purpose, and how the student checked the output.
Third, assessment should include evidence of process. Drafts, outlines, revision history, lab notes, commit logs, short oral checks, or in-class components can tell us far more than a detector score. If a student can explain their argument, defend their choices, and discuss the material intelligently, that matters.
Fourth, AI detector scores should never be used alone. They should trigger review, not judgement. Any institution using these tools should have a written process that explicitly says the score is not proof. Students should be allowed to respond, provide drafts, explain their workflow, and challenge the interpretation.
Finally, educators need time and training. Not just another PDF policy from central administration, written by someone who thinks “ChatGPT” is a cybersecurity incident. Actual training. Actual examples. Actual discussion of what AI can and cannot show.
The real lesson
The great irony of AI detection is that it tries to solve a trust problem by introducing another trust problem.
Students may misuse AI. That is real.
But detectors may misrepresent students. That is also real.
If an institution is serious about fairness, it cannot only worry about false negatives, where AI-generated work slips through. It also has to worry about false confidence, where a tool produces a number and everyone starts behaving as if the number knows more than it does.
The study is useful because it does not merely say “AI detectors are imperfect”. We already knew that. It shows a more specific and more troubling pattern: mixed human/AI writing can be scored in misleadingly inflated ways, while heavily AI-generated writing can be scored in misleadingly deflated ways. That means the detector is not just noisy. It may be noisy in a way that punishes the wrong behaviour and misses the more serious one.
That should concern every teacher, every student, and every institution currently building policy around these systems.
AI has absolutely created a real assessment problem. But if the solution is to let a black-box detector throw percentages at students and call it academic integrity, then we have not solved the problem. We have just automated the suspicion and given it a dashboard.
And in the end, that may be the most AI-era thing imaginable: using a machine we do not fully trust to detect whether someone used a machine we do not fully trust, so that a human can make a decision they pretend was objective.
Welcome to the future of education.
Please upload your essay, your draft history, your AI declaration, your soul, and a screenshot proving you are not a robot.
PS: this article is 100% human written my my, your grandmaster Kim Schulz and with the lovely help of my trusty Language tool (because, beleive it or not, but English is not my primary language). Image is