AI

Do AI Detectors Work? OpenAI Withdrew Its Own

The company with the most to gain from a working detector built one, measured it, and took it down. The numbers it published say why.

A hand writing in a notebook beside an open book.
Photo: Mikhail Nilov / Pexels
Reading mode

If you buy through our links, we may earn a commission. It never affects our verdicts or scores — how that works. As an Amazon Associate I earn from qualifying purchases.

Not well enough to accuse anyone. The strongest evidence is not an opinion about detectors — it is that OpenAI built one, measured it, and took it down.

The notice still sits at the top of its own announcement page: “As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy.”

The numbers it published on the way out

OpenAI stated them plainly while the tool was still live, which is more than most vendors do. On a challenge set of English texts, the classifier “correctly identifies 26% of AI-written text (true positives) as ‘likely AI-written,’ while incorrectly labeling human-written text as AI-written 9% of the time (false positives).”

Sit with both halves.

It missed roughly three out of four. Anyone using AI and wanting to hide it was not caught by this.

It accused about one human in eleven. In a class of thirty, that is between two and three innocent people, every time you run it.

And this was the company that made the model being tested, with every advantage in doing so. Its own framing concedes the general case: “it is impossible to reliably detect all AI-written text.”

The commercial ones have a specific, non-random failure

Vendors still selling detectors will tell you theirs is better. The most useful independent check is a Stanford evaluation of several widely-used detectors, and its finding is not that they are merely inaccurate.

The detectors “consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified.”

That is worse than random error, because it lands on the same people every time. The authors’ explanation is that detectors are effectively measuring how predictable and constrained the writing is — so a person writing carefully in their second language looks, to the machine, like a machine.

The same paper closes the loop on whether the tools even achieve their purpose: “simple prompting strategies can not only mitigate this bias but also effectively bypass GPT detectors.”

So the tool can be defeated by anyone who reads one blog post about defeating it, while continuing to flag people writing honestly in a second language. The authors “caution against their use in evaluative or educational settings”.

Why this is hard in a way it will stay hard

A detector for AI text is not a fingerprint reader. There is no marker left behind — the output is ordinary words in ordinary order, and the same sentence can be produced by a person or a model.

What detectors actually measure is statistical typicality: how unsurprising each word is given the ones before it. That is a real signal, and it points the wrong way for exactly the people who write plainly, follow a taught structure, or use a limited vocabulary. Clear, conventional prose is what a model produces and what a careful student is taught to produce.

Provenance approaches — signing content at the point of creation — are a different and more tractable problem. OpenAI itself said it was “researching more effective provenance techniques for text”. Detection after the fact is the one that does not have a solution in view.

What to actually do

If you are being accused: ask what the false positive rate is and how it was measured on writing like yours. Ask for evidence that is not the score — a draft history, version timestamps, a conversation about your own argument. A detector output is not evidence of anything on its own, and the paper above is a citable reason why.

If you are a teacher or an editor: do not use a score as the basis of an accusation. Use it, at most, as a prompt to look at process, and be aware that acting on it disproportionately penalises non-native speakers — which is a fairness problem before it is an accuracy one.

If you commission writing: ask for drafts and sources rather than running a detector. Wanting to know whether the work is any good and whether the claims check out is a better question, and it is one you can actually answer.

Either way, do not pay for certainty here. The organisation best placed to sell it looked at its own numbers and withdrew the product instead.

How we researched this

No one at bitcritiq has handled this product. Everything here comes from published sources, listed below.

What this cannot tell you
This is about detectors for written text. Image, audio and video provenance is a different problem with different tooling, and some of it works considerably better because it can rely on signatures added at creation. The paper cited is from 2023 and detectors have changed since; what has not changed is the structural problem it identifies, and no vendor has published an independent evaluation that answers it.
How we chose this, and what we did
Why this subject
People search this from both sides — writers accused of using AI, and teachers deciding whether to trust a score. Both are usually offered vendor confidence or forum anecdote. There is better evidence than either: the numbers OpenAI published about its own detector before withdrawing it, and a peer-reviewed evaluation of the commercial ones.
How we looked at it
Took the accuracy figures and the withdrawal notice from OpenAI's own announcement page, which still carries both. Took the bias finding from the Stanford paper that measured several widely-used detectors against writing by native and non-native English speakers, using the authors' own wording for what they found.

What this rests on

Each claim below, and how firmly it is held. Nothing here was measured by bitcritiq — see how we test for why.

  • OpenAI withdrew its own AI-text classifier for low accuracy.

    OfficialNotice on OpenAI's own announcement page, dated 20 July 2023.

  • That classifier identified 26% of AI-written text and wrongly flagged 9% of human-written text.

    OfficialFigures published by OpenAI while the tool was live.

  • Widely-used detectors systematically misclassify non-native English writing as AI-generated.

    CorroboratedPeer-reviewed evaluation by Liang et al. across several commercial detectors.

Sources 2

  1. New AI classifier for indicating AI-written text — OpenAIOfficialaccessed Aug 29, 2026
  2. GPT detectors are biased against non-native English writers — Liang, Yuksekgonul, Mao, Wu and ZouStandards / .govaccessed Aug 29, 2026

read next

Specifications