Evidence check | 28 August 2026

Are AI detectors accurate?

A detector flag is not proof of AI authorship. The useful questions are: which detector, which texts, which threshold, and how many false alarms?

Verdict: a headline accuracy percentage cannot tell you the probability that one flagged essay was AI-written.

ToolGlance editorial analysis. Published 30 May 2026; evidence updated 28 August 2026. Published research and a transparent calculation, not a new detector test.

AI detector false-positive calculator

Even a detector that catches 99% of AI texts and falsely flags only 1% of human texts can produce as many false alarms as correct flags when AI text is rare. The example below assumes 1% AI text in 10,000 documents.

50.0%of flags are correct in this scenario

AI texts correctly flagged
99
Human texts wrongly flagged
99
AI texts missed
1
Human texts not flagged
9,801

Expected counts under your assumptions, not observed people or measured product performance. No writing is uploaded.

Correct share of flags = (AI share × detection rate) ÷ [(AI share × detection rate) + (human share × false-positive rate)]. Rates are proportions in the formula.

Hypothetical detector with 99 percent detection and 1 percent false positives: at 1 percent AI prevalence, half of flags are correct; at 10 percent prevalence, 91.7 percent are correct.
ToolGlance calculation. Both examples use the same detector rates; only the assumed proportion of AI text changes. This is not a vendor ranking.

What the research actually tested

These studies are not interchangeable. Their populations, dates, detectors and metrics differ, so we do not average them into one accuracy score.

SourcePopulation and scopeFindingWhat it does not establish
OpenAI: retired text classifier2023-01-31 | Provider evaluation; historicalEnglish challenge set; retired productReported 26% detection of AI text and 9% false positives on human text. The classifier was withdrawn on 20 July 2023.Not an estimate of present-day detectors or a current product recommendation.
Stanford HAI: non-native English writing2023-05-15 | University report of research91 TOEFL essays; seven detectorsThe reported average false-positive rate for the TOEFL essays was 61.22% across the tested detectors.A specific historical sample, not the rate for every non-native writer or every detector.
RAID: shared detector benchmark (ACL 2024)2024-08 | Peer-reviewed conference paper2024 paper: 6M+ generations, 11 generators, eight domains; 12 detectors evaluatedPerformance varied with unfamiliar generators, decoding choices and adversarial changes.Use the paper's dataset edition; the current repository contains a larger, expanded dataset.
Park, Jeong and Kim: editing style as a confound2026-08-27 | arXiv v1; venue status not independently verified135,389 original/edited document pairs from one academic editing service, 2018-2025; 13 detectorsEditing changed detector scores in different directions. Baseline false-positive analysis used the pre-ChatGPT subset.Do not treat all 135,389 pairs as the baseline false-positive denominator, or this sample as a current commercial-detector ranking.

The newest entry is the 27 August 2026 arXiv submission. Its paired-document design asks whether professional editing changes detector outputs. It is not a survey of all writing, and it does not give a universal chance that a flagged document is AI-generated.

Before trusting a “99% accurate” claim

Illustration of a reader comparing marked-up manuscript drafts beside a laptop.
AI-generated editorial illustration. Not a photograph of a study, product test or research participant.

Check the denominator

Accuracy is the share of all classifications that are correct. Precision is the share of flags that are correct. Recall is the share of AI texts caught. A supplier must specify which one it reports.

Check the population

A result on short English passages does not automatically transfer to edited academic papers, translated text or another language. Ask for the text length, domain and human/AI balance.

Check the operating point

The score threshold determines which texts get flagged. Comparing products at different false-positive rates can make an apparent winner meaningless. Keep the model version and test date alongside the result.

Keep an independent record

For an actual decision, retain source notes, drafts, revision history and the relevant disclosure rules. A numerical flag is a reason to investigate, not a substitute for those records.

Changing the base rate changes the interpretation even when sensitivity and false-positive rate stay fixed. This is ordinary conditional-probability arithmetic. The calculator does not estimate the true base rate in your classroom, publication or organisation: that remains an explicit assumption.

Data, chart and citation

ToolGlance's original chart and hypothetical scenario table are available under CC BY 4.0. The linked papers and source material retain their own rights. Reuse the chart with its assumptions and source note intact.

Image embed code

Questions about AI detector accuracy

Are AI detectors 99% accurate?

There is no universal accuracy figure. A claim needs the detector version, threshold, dataset and test date. Accuracy also depends on the proportion of AI and human texts; a 99% claim is not a 99% probability that a flagged document is AI-written.

Can a human-written essay be flagged as AI?

Yes. That is a false positive. The size of the risk depends on the detector and population. Keep drafts, sources and revision history; a detector score alone does not establish authorship.

Does this calculator check my writing?

No. It calculates expected counts from rates you enter. It does not read text, run a detector or decide whether someone used AI.

Are these ToolGlance benchmark results?

No. The evidence table attributes published findings to their original researchers. The calculator and scenario CSV are ToolGlance arithmetic, not a new detector benchmark.

Does a lower score after editing prove a text is human?

No. A changed score shows that the detector responded differently. Edit for correctness and readability, follow the applicable disclosure policy, and do not use a score as proof of authorship.

Choose a writing tool for the task

Authorship detection, editing and writing assistance are different jobs. Compare capabilities and disclosure requirements before choosing a product.

Correction record

28 August 2026: replaced an unsourced blanket 83-94% accuracy range with source-specific findings and an explicit base-rate calculation. The URL is unchanged.

Corrections: contact ToolGlance. The research register is an editorial synthesis, not a systematic review or an endorsement of a detector.