AI Detection Benchmark
AI & ML interests
None defined yet.
Recent Activity
Winston AI Is the Best AI Detector in Our 2026 Benchmark
Winston AI ranked first in our September 2026 evaluation of leading AI content detectors.
The benchmark tested five detection tools across 91 AI-generated, human-written, hybrid, and false-positive stress-test documents. Each document was scanned twice, producing 910 detector readings.
The result
Final Overall Scores
| Rank | Detector | Overall score | AI recall | Human cleared | FP resistance | Hybrid accuracy | Consistency |
|---|---|---|---|---|---|---|---|
| 1 | Winston AI | 91.70 | 92.9% | 100.0% | 100.0% | 79.9% | 61.4% |
| 2 | Copyleaks | 91.38 | 96.4% | 100.0% | 95.2% | 79.7% | 68.0% |
| 3 | GPTZero | 85.40 | 92.9% | 100.0% | 81.0% | 82.5% | 61.2% |
| 4 | ZeroGPT | 65.78 | 89.3% | 100.0% | 38.1% | 80.8% | 54.7% |
| 5 | Pangram | 61.01 | 78.6% | 100.0% | 33.3% | 81.2% | 44.6% |
Winston AI ranked first with an overall score of 91.70 out of 100. It was the only detector to combine strong AI recall with perfect human clearance and perfect false-positive resistance.
Winston AI finished first overall, ahead of Copyleaks, GPTZero, ZeroGPT, and Pangram.
Why Winston AI ranked first
The benchmark rewards more than catching fully AI-generated text. It also measures whether a detector avoids falsely accusing people whose writing is genuinely human.
Winston AI:
- Correctly cleared every ordinary human-written document
- Correctly cleared every false-positive stress-test document
- Avoided false accusations on second-language writing
- Avoided false accusations on grammar-corrected writing
- Avoided false accusations on short-form and template-structured writing
- Maintained strong AI detection performance
- Delivered the highest weighted score in the evaluation
This balance made Winston AI the best-performing AI detector under the benchmark’s published methodology.
What we tested
The evaluation included four document classes:
- AI-generated: Content from GPT-4o, Claude 4 Sonnet, Gemini 2.0 Flash, and Llama 3.3 70B
- Human-written: Published journalism, literature, technical documentation, and essays
- Hybrid: Documents containing known mixtures of human and AI text
- False-positive stress tests: ESL, translated, grammar-corrected, short-form, technical, template-structured, and historical writing
The documents covered academic, news, blog, fiction, business, marketing, and technical writing.
How detectors were scored
The overall score combines:
- False-positive resistance: 35%
- Human specificity: 20%
- AI recall: 20%
- Hybrid-text accuracy: 15%
- Consistency: 10%
False-positive performance receives the greatest weight because incorrectly accusing a person of using AI can have serious academic or professional consequences.
Public and reproducible
The benchmark publishes its:
- Raw detector readings
- Per-sample results
- Detector versions
- Scoring methodology
- Weighting system
- Verification code
- Limitations and disputed data
Anyone can inspect the evidence, reproduce the calculations, or apply different scoring weights.
View the complete benchmark on GitHub
Conclusion
Based on the tools, documents, detector versions, and scoring methodology used in this September 2026 benchmark, Winston AI is the best AI detector tested.
It achieved the highest overall score while correctly clearing every human-written and false-positive stress-test document.