Pangram 4 Technical Report
2026-07-29 • Computation and Language
Computation and Language
AI summaryⓘ
The authors present Pangram 4, a new AI model designed to tell if text is written by humans, AI, or a mix of both. Their model is very accurate, making few mistakes and working well even on texts that it hasn't seen before or that try to trick it. They also improved the model's ability to spot small changes and mixed authorship in text. Overall, Pangram 4 performs better than earlier versions and other similar models on standard tests for detecting AI-generated text.
AI text classificationAUROCfalse positive ratefalse negative rateout-of-distribution generalizationadversarial attacksboundary detectionhuman-AI co-authorshipAI text detection benchmarks
Authors
Ben Glickenhaus, Katherine Thai, Jenna Russell, Elyas Masrour, Yue Han, Max Spero, Bradley Emi
Abstract
We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, Pangram 4 exhibits superior out-of-distribution generalization and robustness to adversarial attacks. Another novel contribution of Pangram 4 is its improved ability to distinguish fine-grained edits and mixed AI-human co-authored text. We demonstrate improvements to both boundary detection tasks and the detection of interleaved AI assistance. Finally, we report metrics on standard AI detection benchmarks showing that Pangram 4 achieves state-of-the-art performance on the AI text detection task across a wide variety of settings and domains.