The company that built the AI detector dropped it first
Parents ask whether to run a draft through a detector. What those tools actually measure is not who wrote the essay — it is how predictable the vocabulary is. Which is why the honest writer is the one who gets caught.
It comes up in almost every conversation: should we run the essay through an AI detector? We understand the impulse. Here is the uncomfortable part.
The company that built one dropped it first
OpenAI, which made ChatGPT, released its own detector and quietly withdrew it six months later. The stated reason was one line: it was not accurate enough.
How inaccurate: of 100 machine-written texts it caught 26. And in the other direction — of 100 texts written by people, it flagged 9 as machine-written.
What one percent looks like at scale
Turnitin, widely used in schools, put its false-positive rate at 1%. One percent sounds tolerable. Vanderbilt ran its own numbers: roughly 75,000 papers had gone through Turnitin in 2022. One percent of that is about 750 papers a year wrongly flagged.
Vanderbilt disabled the feature in August 2023.
In April 2025 the FTC acted against a detector advertised as 98% accurate; on general-purpose writing the measured figure was 53% — a coin flip. The 98% had come from a study on academic text, and was carried over into advertising for everything else.
The errors are not evenly distributed
A Stanford team ran a simple experiment. They assembled two sets of essays, all written by people: one from American eighth-graders, one from TOEFL writers whose first language was not English.
Through seven widely used detectors, the American students' essays passed as human almost every time. The non-native writers' essays — all genuinely human — were flagged as machine-written roughly six times in ten.
Why: a detector does not look at who wrote the text. It looks at how predictable the next word is. The safer and more common the vocabulary, the more machine-like it scores.
The tool is not measuring authorship. It is measuring vocabulary range.
There is a bitter turn. The researchers took the falsely flagged essays and asked ChatGPT to polish them. The verdict flipped to human. A student who knows how to launder a draft through AI passes; a student who wrote it honestly does not.
"But my child doesn't use those tools"
That is exactly why it matters. The Max Planck Institute analysed roughly 740,000 hours of people actually speaking — not writing. After ChatGPT's release, the words the model favours rose sharply in ordinary human speech. Nobody asked anyone to do that.
So an essay that reads like a machine is not evidence of cheating. It is this generation's default.
And the machine is arriving from the other side too
In a 2023 survey, eight in ten higher-education institutions said they planned to bring an AI tool into review within the year. In 2025 Virginia Tech changed its model outright: essays that used to be read by two humans are now read by one human and one AI. If the two scores differ by more than two points a second human is brought in, and the final decision stays with admissions staff.
Which leaves an honest student between two blades — language drifting toward the machine on one side, and a machine trying to filter for that on the other.
There is something that works. It is not a machine.
A UMass team gave five readers who use AI writing tools daily a mixed set of 300 human- and machine-written articles. The majority vote of those five got all but one of the 300 right. Commercial detectors, on the same material, sat around 85%.
The part that matters more: those readers could say what they had noticed. Sentences that are only ever tidy. Grammar that is suspiciously flawless. An ending that dissolves into vague optimism. Not a hunch — a skill that can be taught.
Which is the work we do
We hold over 800 real college application essays and about 700 written evaluations in which experienced reviewers said why a given essay earned its rating. That is the instrument. Not instinct — counts.
With an instrument the advice changes. Instead of "maybe add some dialogue," we can say that top-rated essays contain real remembered dialogue about six times in ten, and the bottom of the range about one time in ten. That is evidence rather than exhortation.
From those essays we drew fourteen criteria. One question runs through all of them.
Swap in a different applicant, change only the name — does the essay still stand?
If it does, it does not yet belong to anyone.
What we do not give
- A score
- No.
- A probability
- No.
- A verdict that a machine wrote it
- No. We do not adjudicate authorship.
- A comparison against a cohort average
- No. The only comparison is the student's own previous draft.
A detection score asks whether a machine wrote it. An admissions officer asks the better question — could anyone else have written this? We coach for the second one, because once the answer to it is no, the first question stops mattering.
SourceOpenAI's own announcement · Vanderbilt (Aug 2023) · FTC v. Workado (Apr 2025) · Stanford, Liang et al., Patterns 2023 · Max Planck · UMass, Russell et al., ACL 2025 · Intelligent survey 2023 · Virginia Tech 2025. Citations re-checked August 2026. The reading built on them is IPE's own.