Skip to main content
GlyphFlow
Writing Tips Content Strategy AI Writing SEO

Why AI Writing Detectors Get It Wrong

Team GlyphFlow

TL;DR

  • Most “AI detector” tools score text using signals like sentence-length uniformity, which sound scientific but are weak and easy to fake.
  • One signal holds up better under peer review: a specific set of words (delve, underscore, meticulous, boast, tapestry) increased sharply in academic writing after ChatGPT’s release, by hundreds to well over a thousand percent in one multi-database study, with independent research confirming the same words spiked.
  • A Stanford study found commercial AI detectors misclassified about 61% of non-native English essays as AI-generated, versus roughly 5% of the native-English comparison set, using the exact same tools.
  • Real students have faced discipline, failing grades, and lawsuits over detector false positives. GlyphFlow doesn’t ship an AI detector, and this is why.

Type “AI detector” into Google and a dozen tools will hand you a confident-sounding percentage: 87% AI-generated, 12% human. The research behind these tools tells a much messier story, and the gap between what a detector claims and what it can actually prove has already cost students real consequences.

The One Signal Backed by Real Research

Most AI-detection tools lean on stylistic signals like “burstiness,” the idea that human writing varies sentence length more than AI writing does. It’s intuitive, and it’s also trivially defeated. Ask any model to vary its sentence length and it will, instantly.

This is a different problem from watermarking, which only works on content a model chose to mark at creation time. A detector claiming to spot unmarked AI text from style alone is making a much bigger claim, and has to earn it.

One signal has held up better, mostly because it wasn’t designed as a detection method at all. Researchers tracking word frequency across millions of academic papers noticed that a specific, narrow set of words started appearing far more often once ChatGPT became widely used. A peer-reviewed analysis of PubMed abstracts tested 135 candidate words against 84 control phrases and found 103 of them crossed a statistical threshold for a meaningful increase, “delve,” “underscore,” “meticulous,” and “boast” among the strongest. A separate multi-database study covering Scopus, PubMed, and several other indexes independently found the same words spiking, with “delve” and “underscore” among the largest movers between 2022 and 2024, though the exact percentages differ by database and by whether the count is titles-and-abstracts or full text. The two studies don’t agree on a single number, which is itself informative: the effect is real and reproducible, but its exact size depends heavily on how you measure it, and neither paper claims the increase alone proves any single document was AI-written.

A third analysis, testing several open-source models directly, found the same word list and tested explanations for why it happens. Training-data overrepresentation didn’t hold up as the cause. The evidence pointed instead toward the fine-tuning process, specifically the human-feedback stage where reviewers rate model outputs. The researchers’ theory: rushed human raters may lean on the presence of certain words as a shortcut for judging quality, and the model learns to produce them regardless of whether the content actually warrants them.

This is a real, reproducible pattern with a plausible, tested explanation behind it. It’s also a moving target. Once a word gets flagged as an AI tell, it’s the first thing a “make this sound more human” pass removes.

Why “Does It Sound Robotic” Fails

Vocabulary drift is measurable and has a documented cause. Most of what detectors actually run on doesn’t hold up nearly as well.

A 2023 Stanford study published in Patterns ran seven commercial AI detectors against two groups of essays: a native-English comparison set (US eighth-grade student essays) and a set of TOEFL exam essays written by non-native English speakers. The native set was misclassified as AI-generated about 5.19% of the time. The TOEFL set was misclassified 61.22% of the time, by the same seven tools. The comparison isn’t perfectly matched (eighth-graders versus adult test-takers), which the paper itself acknowledges, but it points to a specific, tested cause rather than a hunch: the researchers found that even simplifying the vocabulary in the native essays pushed their misclassification rate up to 56.65%, which suggests the detectors are penalizing plain, less varied word choice generally, not non-native writing specifically. That’s simply how someone writes in a language they’re still building fluency in, and it has nothing to do with who or what wrote the piece.

The Cost of Getting It Wrong

This isn’t an abstract accuracy problem. It has already produced formal discipline and federal lawsuits, and not always with a detector tool even in the picture.

At the University of Michigan, an undergraduate sued the university in federal court after an instructor accused her of using AI across three papers. The lawsuit doesn’t describe any commercial AI-detection software being used. Instead, the instructor’s evidence was his own comparison output, generated by prompting an AI system with the student’s outline and his own feedback, then treating the resemblance as proof. Her lawsuit says her documented anxiety and OCD produce the same traits he flagged as suspicious: formal tone, meticulous structure, and stylistic consistency. She received disciplinary sanctions and a “no record” grade before the case reached a courtroom, and the case remains active litigation as of this writing.

At UC Davis, a professor told William Quarterman “it appears as though this exam is plagiarized” and that one answer was “drawn from ChatGPT or similar AI software,” based on GPTZero’s output. She gave him a failing grade and referred him for academic dishonesty. He appealed with a Google Docs edit history showing his actual writing process, along with research on GPTZero’s unreliability, and was eventually cleared.

At Yale, an executive-MBA student named Thierry Rignol was suspended for a year and given a failing grade in the course after GPTZero flagged his final exam. Rignol, a French national, argues the accusation reflects the same non-native-English bias the Stanford research documented. His federal lawsuit, Rignol v. Yale, is still active; a judge already denied his request for a preliminary injunction in May 2025, finding he hadn’t shown the suspension caused irreparable harm. Yale disputes his account and points to evidence beyond the detector score, including questions about the source file behind his submission.

None of these cases is a clean, one-cause story, and two of the three are still working through the courts. What they share is a pattern: something that read as unusual, whether to an instructor’s ear or a detector’s score, got treated as proof, and the people paying the price were disproportionately the ones whose writing looked different for reasons that had nothing to do with AI.

What This Means If You’re Checking Someone Else’s Writing

AI-assisted writing does happen, and the vocabulary research above is genuine evidence of it at scale. It is not, on its own, a basis for a confident percentage score about one specific piece of writing. Turning population-level word-frequency data into a verdict about a single essay is a bigger claim than the underlying research supports.

A more defensible approach looks less like running text through a black-box scorer and more like what an experienced teacher already does by instinct: compare it against that person’s own earlier work. Does the vocabulary, structure, and rhythm match how they’ve written before? That’s a question about consistency, not a verdict about authorship, and it doesn’t carry the same risk of penalizing someone for writing carefully, or for writing in a second language.

It’s the same instinct behind writing genuinely helpful, people-first content instead of chasing a metric: judge the actual thing in front of you instead of a proxy for it. GlyphFlow doesn’t ship an AI-detection score, and after this research, we don’t plan to. The signals underneath every existing tool are either too weak to trust or too new to have been stress-tested against the population most likely to be harmed by a false positive. If you’re trying to understand your own writing rather than judge someone else’s, our readability tools measure sentence structure and clarity directly, without claiming to know who, or what, wrote it.

People Also Ask

Can AI detectors be trusted for grading or academic decisions? Independent research has found major accuracy gaps, particularly against non-native English writers, who were misclassified at more than ten times the rate of a native-English comparison group in a peer-reviewed Stanford study. A number of universities, including Vanderbilt, have disabled or stopped relying on commercial AI detectors for disciplinary decisions as a result.

What’s the most reliable sign of AI-generated text? A documented shift in specific vocabulary, words like “delve,” “underscore,” “meticulous,” and “boast,” which increased sharply in academic writing after ChatGPT’s release, according to multiple peer-reviewed word-frequency studies. It’s a real, measurable pattern at the population level, though not one any single piece of writing can be judged against with certainty.

Why do AI detectors flag human writing as AI-generated? Most detectors score text using stylistic patterns like sentence-length uniformity or plain, low-variety word choice, both of which occur naturally in writing by non-native English speakers and anyone whose personal style happens to be consistent. The tool can’t tell “written by AI” apart from “written carefully, in a smaller vocabulary.”

My Take

The vocabulary research here is genuinely interesting, and it’s one of the few things in this space backed by numbers that held up under peer review, even if different studies land on different exact figures. But research and a trustworthy product are different things. The moment you turn “these words got more common across millions of papers” into “this one essay is 74% AI-generated,” you’ve made a claim the underlying research never supported, and put that claim in a position to end someone’s semester. A Michigan student is now suing over an accusation built on nothing more than an instructor’s own prompt. A Yale student is suing over a detector known to misfire on non-native English writing, his own. A UC Davis student had to produce his entire drafting history to undo a failing grade a single tool assigned him. Until detection methods can account for the writers they’re most likely to misjudge, a percentage score is a confidence trick, not a measurement.

Recommended Writing Guides

AI Writing

Anthropic Is Now Watermarking Claude's Text

Anthropic now embeds invisible watermarks in Claude's text. Here's how it works, who can detect it, and what it means if you write with AI.

By Team GlyphFlow
Writing Tips

Can AI Influencer Networks Really Work?

AI tools now let brands clone a creator's style and voice into a synthetic avatar at scale. We dissect the legal and ethical questions this raises.

By Team GlyphFlow