What AI detectors actually measure.

Your page got flagged. The tool that flagged it cannot tell you who wrote anything, and the number it produced is not the number that decides whether an answer engine ever mentions you.

Mike Millett July 21, 2026 8 minute read

It usually arrives as a screenshot. A percentage, a red bar, and a question underneath it: is this bad.

Someone has run a page through an AI detector. The page came back 87% AI, or 94%, or 61%, and now there is a conversation happening about whether the writer lied, whether the agency cheated, whether the site is about to be punished.

That number is real in the sense that a real tool produced it. It is not evidence of anything, and I can show you why with sources you can check yourself.

The case detectors handle well is a rounding error

Start with what is actually out there. In April 2025 Ahrefs sampled 900,000 newly published English pages, each from a different domain, and sorted them by how they were written.

2.5%

Share of new web pages that are purely AI written, with no human editing. Another 25.8% were purely human. The remaining 71.7% were a blend of the two.

Ahrefs, 900,000 newly discovered English pages, one per domain, April 2025.

Detection is genuinely tractable on that 2.5%: long, unedited, straight-from-the-model output has a recognizable statistical signature. But almost nothing on a working business website is that. Your pages were drafted, edited, argued over, pasted from an older version, half rewritten by the owner on a Sunday. They live in the 71.7%, which is exactly where the tools stop working.

One caveat, against my own source

Ahrefs produced those proportions using its own AI detector, which is the class of instrument I am about to argue against. So take the shape of the finding, not the decimal places. The shape is not seriously disputed: pure machine output is a small slice, and mixed authorship is the norm.

On everything else, the tools disagree with themselves

Eight academic integrity researchers ran the most comprehensive public test of this to date. Weber-Wulff and colleagues evaluated 14 detection tools, including the two commercial systems most schools actually pay for, Turnitin and PlagiarismCheck.

Not one tool reached 80% accuracy. Only five cleared 70%. Light obfuscation, the kind any editor performs by accident, degraded performance further. Their published conclusion is one sentence long: the available detection tools are "neither accurate nor reliable."

Where that study cuts against me

Their tools' dominant error was the opposite of the one people fear. The systems leaned toward calling text human written, which means missed AI rather than falsely accused humans. On its own, that study is not evidence that your page was wrongly flagged. For that, you need the next one.

The experiment that should end the argument

In 2023, a Stanford group published a study in Patterns that ran seven widely used detectors against two sets of human writing: 91 TOEFL essays written by non-native English speakers, and 88 US eighth-grade essays.

On the eighth graders the detectors were close to perfect. On the TOEFL essays they were catastrophic. The average false positive rate was above 61%. All seven detectors unanimously declared 18 of the 91 essays to be machine written. Across the set, 89 of 91 were flagged by at least one tool.

Every one of those essays was written by a person.

Then the researchers ran the experiment in reverse, and this is the part that matters. They took the native-written eighth-grade essays, the ones the detectors had cleared, and asked ChatGPT to change one thing: simplify the word choices, as if written by a non-native speaker. Same authors. Same arguments. Plainer vocabulary.

5.19% to 56.65%

False positive rate on human-written essays before and after their word choices were simplified. Nothing about the authorship changed. Only the vocabulary did.

Liang, Yuksekgonul, Mao, Wu and Zou, Patterns, Cell Press, July 2023.

The detector was never tracking who wrote the text. It was tracking how plainly the text was written, and reporting that as authorship.

What that costs, when someone believes it

Common Sense Media surveyed 1,045 US teenagers between March and May 2024. About one in ten said work they had written themselves was wrongly identified as AI generated. The rate was not evenly distributed: 20% of Black teens reported a false accusation, against 10% of Latino teens and 7% of white teens. Most of the students who were wrongly flagged had their work run through detection software.

How much weight this carries

That is self-reported survey data from teenagers, not adjudicated cases, and I am not going to present it as more than that. It is corroborating evidence, not proof. What makes it worth including is that it points in the same direction as the Stanford result, from a completely different method: the burden of the error lands hardest on people whose English the machine finds unusual.

The company that made the text could not find the text

OpenAI shipped its own AI text classifier on 31 January 2023. At launch it correctly identified 26% of AI-written text, and wrongly flagged human writing as AI 9% of the time.

26%

Share of AI-written text OpenAI's own detector caught. It was withdrawn on 20 July 2023, in the company's words, "due to its low rate of accuracy."

OpenAI, AI Text Classifier launch post and withdrawal notice, 2023.

The organization with the most access to the model, the most technical capability, and the most commercial reason to solve this, built a detector that missed three quarters of its own output and retired it in under six months. Every vendor still selling you a confident percentage is claiming to have solved a problem that OpenAI publicly gave up on.

This is a ceiling, not a product defect

There is a reason the tools keep failing in the same direction, and it is not that nobody has built a good one yet.

Sadasivan and colleagues proved the bound. The performance of any detector is limited by how much the machine's text distribution overlaps with human text. Written formally, the area under the ROC curve of the best possible detector is at most one half plus the total variation distance between the two distributions, minus half that distance squared. Written plainly: as language models get better at producing text that looks like ours, the distributions converge, and the ceiling on detection accuracy falls toward a coin flip.

Not the ceiling on today's detectors. The ceiling on all of them, including the ones not built yet. Every improvement in generation is, by construction, a degradation in detectability. That is the trade the field is stuck with.

The strongest version of the other side

Turnitin claims a document-level false positive rate below 1%, and from outside the company I cannot disprove it. Detectors do beat chance on long, pure, unedited output. So the honest claim is narrower than "detectors are wrong."

It is also worse. Vanderbilt did the arithmetic on that 1% and disabled the feature in August 2023: the university had submitted 75,000 papers the previous year, so a 1% error rate means roughly 750 students wrongly implicated. Their stated reason was that Turnitin "gives no detailed information as to how it determines if a piece of writing is AI-generated." A number you cannot audit, applied at scale, produces casualties even when the number is good.

So what is the score actually measuring

Two things, mostly: perplexity, meaning how predictable each next word is, and burstiness, meaning how much sentence length varies. Text that is evenly paced, plainly worded, formal in register and low in surprise sits in the region the model associates with machine output.

Now read that list again and picture what it describes on a business website. Compliance copy. Policy summaries. Service descriptions. Safety notices. Anything written carefully and plainly by a person trying not to confuse a nervous customer. The clearer your writing, the more machine-like it scores.

It is closer to a plainness detector than an authorship detector. Which would be a harmless curiosity, except that people are making decisions with it.

Meanwhile, the thing you were actually worried about

Underneath the panic is usually a business fear rather than a moral one: that Google is going to find out and punish the site.

Google does not run an AI detector against your pages, and there is no blanket penalty for AI-assisted writing. The policy that does exist is scaled content abuse, formalized in March 2024: generating many pages primarily to manipulate rankings, with little or no value added for readers. It is deliberately method-agnostic. A thousand thin pages written by a human violate it exactly as much as a thousand written by a model.

One update is worth knowing. On 15 May 2026 Google widened its spam definition to cover "attempting to manipulate generative AI responses in Google Search," bringing AI Overviews and AI Mode explicitly inside the policy.

So the risk was never that a machine helped you write. The risk is publishing thin, unoriginal, unsourced pages at volume. And unlike authorship, that is measurable without guessing.

The question that actually pays

Replace "will this get flagged" with a question that has an answer: does this page give an answer engine a reason to cite it?

That one is decidable from your own HTML, deterministically, same input and same result every time. Roughly seven things carry it:

Is the text in the raw HTML at all. AI crawlers do not execute your JavaScript. If the words only appear after a script runs, an answer engine reads an empty page.

Is a real, named person credited in a way a machine can parse, rather than a byline buried in prose that a crawler cannot distinguish from a testimonial signature.

Do the claims link to primary sources, quote named people, and put statistics next to where they came from.

Is anything here first-hand. Your photographs, your data, your table. The part a text generator cannot invent on your behalf.

Is it specific. Named places, real prices, actual dates, concrete numbers. Specificity is the exact inverse of generic, and it is what gets quoted.

Does it read as templated, which is the one signal that overlaps with what detectors measure, reported honestly as a style observation rather than an accusation.

Are the dates honest, or has the page been stamped fresh without the words changing.

Not one of those requires knowing who typed the page. All of them decide whether an assistant quotes you or the competitor two miles away.

What we built, and what it refuses to do

We put this into a free check, currently in beta. It scores any page you own against those seven pillars, fetching it exactly as an AI crawler receives it, and hands back the specific defects with the evidence attached.

It will never tell you your content was written by AI, because nobody can honestly tell you that. It will tell you whether your writing carries the surface patterns detectors react to, and it will show you the sentences that do, which is a claim we can actually defend.

If we sold you a percentage instead, the first client who asked us to prove it would be owed an apology.

Two questions worth asking any vendor who hands you one. What is your false positive rate on plain English written by a human. And what happens to that rate when the writer is not a native speaker.

The tool cannot know who wrote your page. It is telling you how you write. That is genuinely worth knowing, and it is a completely different sentence.

Sources

  1. Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou, "GPT detectors are biased against non-native English writers," Patterns, Cell Press, July 10, 2023. Seven detectors, 91 TOEFL essays by non-native writers, 88 US eighth-grade essays. Source of the 61% average false positive rate, the 18 unanimous misclassifications, and the 5.19% to 56.65% simplification result. cell.com/patterns
  2. Debora Weber-Wulff, Alla Anohina-Naumeca, Sonja Bjelobaba, Tomáš Foltýnek, Jean Guerrero-Dib, Olumide Popoola, Petr Šigut and Lorna Waddington, "Testing of detection tools for AI-generated text," International Journal for Educational Integrity 19(1), article 26, December 2023. Fourteen tools including Turnitin and PlagiarismCheck. doi.org/10.1007/s40979-023-00146-z
  3. Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang and Soheil Feizi, "Can AI-Generated Text be Reliably Detected?" arXiv:2303.11156, March 2023. Source of the impossibility bound on detector AUROC as human and model text distributions converge. arxiv.org/abs/2303.11156
  4. OpenAI, "New AI classifier for indicating AI-written text," January 31, 2023, and its subsequent withdrawal notice of July 20, 2023, citing "its low rate of accuracy." Launch figures: 26% of AI-written text correctly identified, 9% of human text falsely flagged. openai.com
  5. Ahrefs, "What percentage of new content is AI-generated?" Study of 900,000 newly discovered English pages, one per domain, April 2025. Pure human 25.8%, mixed 71.7%, pure AI 2.5%. Measured with Ahrefs' own in-house detector. ahrefs.com
  6. Common Sense Media, teen and parent survey on generative AI, fielded March to May 2024, 1,045 US teenagers aged 13 to 18. Source of the false-accusation rates of 20%, 10% and 7%. Self-reported. Education Week coverage
  7. Vanderbilt University, "Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector," August 16, 2023. Source of the 75,000 papers arithmetic and the transparency objection. vanderbilt.edu
  8. Google Search Central, "Google Search's guidance about AI-generated content," and the spam policy on scaled content abuse formalized March 2024. developers.google.com
  9. Search Engine Land, "Google updates search spam policies to clarify it applies to generative AI responses," May 15, 2026. Source of the widened spam definition covering attempts to manipulate generative AI responses in Search. searchengineland.com

Trust First

Ask the answerable question.

Nobody can tell you who wrote your page. What can be measured is whether an answer engine has a reason to cite it. That check is free, and it shows its work.