Evidence check | 28 August 2026
Are AI detectors accurate?
A detector flag is not proof of AI authorship. The useful questions are: which detector, which texts, which threshold, and how many false alarms?
Verdict: a headline accuracy percentage cannot tell you the probability that one flagged essay was AI-written.
ToolGlance editorial analysis. Published 30 May 2026; evidence updated 28 August 2026. Published research and a transparent calculation, not a new detector test.
AI detector false-positive calculator
Even a detector that catches 99% of AI texts and falsely flags only 1% of human texts can produce as many false alarms as correct flags when AI text is rare. The example below assumes 1% AI text in 10,000 documents.
of flags are correct in this scenario
- AI texts correctly flagged
- 99
- Human texts wrongly flagged
- 99
- AI texts missed
- 1
- Human texts not flagged
- 9,801
Expected counts under your assumptions, not observed people or measured product performance. No writing is uploaded.
Correct share of flags = (AI share × detection rate) ÷ [(AI share × detection rate) + (human share × false-positive rate)]. Rates are proportions in the formula.

What the research actually tested
These studies are not interchangeable. Their populations, dates, detectors and metrics differ, so we do not average them into one accuracy score.
| Source | Population and scope | Finding | What it does not establish |
|---|---|---|---|
| OpenAI: retired text classifier2023-01-31 | Provider evaluation; historical | English challenge set; retired product | Reported 26% detection of AI text and 9% false positives on human text. The classifier was withdrawn on 20 July 2023. | Not an estimate of present-day detectors or a current product recommendation. |
| Stanford HAI: non-native English writing2023-05-15 | University report of research | 91 TOEFL essays; seven detectors | The reported average false-positive rate for the TOEFL essays was 61.22% across the tested detectors. | A specific historical sample, not the rate for every non-native writer or every detector. |
| RAID: shared detector benchmark (ACL 2024)2024-08 | Peer-reviewed conference paper | 2024 paper: 6M+ generations, 11 generators, eight domains; 12 detectors evaluated | Performance varied with unfamiliar generators, decoding choices and adversarial changes. | Use the paper's dataset edition; the current repository contains a larger, expanded dataset. |
| Park, Jeong and Kim: editing style as a confound2026-08-27 | arXiv v1; venue status not independently verified | 135,389 original/edited document pairs from one academic editing service, 2018-2025; 13 detectors | Editing changed detector scores in different directions. Baseline false-positive analysis used the pre-ChatGPT subset. | Do not treat all 135,389 pairs as the baseline false-positive denominator, or this sample as a current commercial-detector ranking. |
The newest entry is the 27 August 2026 arXiv submission. Its paired-document design asks whether professional editing changes detector outputs. It is not a survey of all writing, and it does not give a universal chance that a flagged document is AI-generated.
Before trusting a “99% accurate” claim

Check the denominator
Accuracy is the share of all classifications that are correct. Precision is the share of flags that are correct. Recall is the share of AI texts caught. A supplier must specify which one it reports.
Check the population
A result on short English passages does not automatically transfer to edited academic papers, translated text or another language. Ask for the text length, domain and human/AI balance.
Check the operating point
The score threshold determines which texts get flagged. Comparing products at different false-positive rates can make an apparent winner meaningless. Keep the model version and test date alongside the result.
Keep an independent record
For an actual decision, retain source notes, drafts, revision history and the relevant disclosure rules. A numerical flag is a reason to investigate, not a substitute for those records.
Changing the base rate changes the interpretation even when sensitivity and false-positive rate stay fixed. This is ordinary conditional-probability arithmetic. The calculator does not estimate the true base rate in your classroom, publication or organisation: that remains an explicit assumption.
Data, chart and citation
ToolGlance's original chart and hypothetical scenario table are available under CC BY 4.0. The linked papers and source material retain their own rights. Reuse the chart with its assumptions and source note intact.
Questions about AI detector accuracy
Are AI detectors 99% accurate?
There is no universal accuracy figure. A claim needs the detector version, threshold, dataset and test date. Accuracy also depends on the proportion of AI and human texts; a 99% claim is not a 99% probability that a flagged document is AI-written.
Can a human-written essay be flagged as AI?
Yes. That is a false positive. The size of the risk depends on the detector and population. Keep drafts, sources and revision history; a detector score alone does not establish authorship.
Does this calculator check my writing?
No. It calculates expected counts from rates you enter. It does not read text, run a detector or decide whether someone used AI.
Are these ToolGlance benchmark results?
No. The evidence table attributes published findings to their original researchers. The calculator and scenario CSV are ToolGlance arithmetic, not a new detector benchmark.
Does a lower score after editing prove a text is human?
No. A changed score shows that the detector responded differently. Edit for correctness and readability, follow the applicable disclosure policy, and do not use a score as proof of authorship.
Choose a writing tool for the task
Authorship detection, editing and writing assistance are different jobs. Compare capabilities and disclosure requirements before choosing a product.
Correction record
28 August 2026: replaced an unsourced blanket 83-94% accuracy range with source-specific findings and an explicit base-rate calculation. The URL is unchanged.
Corrections: contact ToolGlance. The research register is an editorial synthesis, not a systematic review or an endorsement of a detector.