A new AI text detector claims a drastic reduction in false positives, achieving a claimed error rate of just one mistake per 24,000 documents. Israeli startup Pangram Labs launched the tool on March 26, aiming to solve the core problem plaguing existing detectors: the frequent mislabeling of human-written text as AI-generated.
The tool specifically targets the “false positive” crisis that has led to academic penalties, content suppression, and brand damage. Pangram states its proprietary model is trained on trillions of tokens across formal and informal writing styles, allowing it to distinguish between human variability and machine patterns with higher accuracy.
The False Positive Problem
Existing AI detectors often struggle with non-native English speakers, technical jargon, and highly structured writing. This has resulted in students being wrongly accused of cheating, SEO content being rejected, and businesses losing trust in automated moderation.
Pangram’s approach prioritizes precision over recall. The company suggests that while other detectors catch most AI text, they also flag legitimate human work. Pangram claims its detector is tuned to minimize that risk, with the 1-in-24,000 statistic cited as a benchmark for reliability in high-stakes environments.
“Our model is specifically tuned to minimize false positives,” a Pangram representative said. “We believe accuracy at scale is the only way to make detection useful for publishers, schools, and platforms.”
How the Detector Works
The detection model analyzes linguistic pattern repetition, syntactic uniformity, and probability distributions unique to large language models. It compares text against a baseline of human variability across different domains, such as academic papers, blog posts, and social media comments.
Key differentiators include:
- Training scope: The model was trained on “trillions of tokens” spanning formal and informal contexts, including code and conversational text.
- Granular scoring: Outputs include a confidence score, not a binary “AI/human” verdict.
- Flexible thresholds: Users can adjust sensitivity to prioritize either catching AI text (high recall) or protecting human writing (high precision).
Limitations and Open Questions
Pangram did not release independent third-party benchmarks or a public test dataset. The 1-in-24,000 figure is the company’s internal claim, not a verified standard.
Many existing detectors degrade when tested on adversarial prompts, heavily paraphrased text, or short snippets. Pangram has not clarified how its model performs under those conditions.
The tool is currently available via API and a web dashboard. Pricing and full technical documentation have not been published. The company has also not disclosed whether the model can detect AI writing from older models, such as GPT-3, or newer systems like GPT-4o and Claude 3.
Context for the Market
The AI detection market is crowded but troubled. Major platforms like Turnitin and OpenAI’s own classifier have both admitted to high false positive rates, leading to widespread user frustration. Pangram enters this space with a focused promise: reduce errors to the point where detection is usable for moderation and grading.
If the claim holds, it could reshape how schools, newsrooms, and social platforms handle automated content moderation. If not, it risks joining the long list of tools that over-promise and under-deliver on AI detection accuracy.
Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.
What are your thoughts on this? I’d love to hear about your own experiences in the comments below.