Pangramโs Spero finds AI detection fails against GPT-4
Max Spero of Pangram found that AI text detection accuracy drops over 20% with newer models and is now below 50% for tools tested against GPT-4 and similar systems. He warns that without better verifโฆ
Max Spero, the chief AI officer at Pangram, warned that detecting AIโgenerated text is becoming harder than telling whether a story is real or fake, saying the task is โa moving targetโ in a tweet on Thursday. He said Pangramโs internal research shows that the accuracy of existing detection tools drops by more than 20โฏ% when applied to the latest large language models, and that the gap is widening each month. Spero added that the industry must rethink its approach to verification if it wants to keep up with the pace of AI development.
The problem is rooted in the rapid evolution of language models. As models grow larger and incorporate more diverse training data, they produce text that mimics human style with increasing fidelity. Detection systems that rely on statistical fingerprints or tokenโfrequency patterns are struggling to keep up, because the differences between human and machine output are shrinking. Spero explained that earlier models left obvious tellโtale signalsโsuch as repetitive phrasing or unnatural syntaxโthat detectors could flag. Newer models, however, generate more varied and contextโaware prose, blurring the line between human and AI authorship. The result is a โcat and mouseโ game in which detection tools lag behind every major release.
Pangramโs research team has tested over a dozen detection frameworks against a benchmark set of texts generated by GPTโ4, Claudeโ2, and Geminiโ1. The results were sobering: only 48โฏ% of samples were correctly identified as AIโgenerated by the bestโperforming tool, and accuracy fell to 35โฏ% for the most recent models. Spero highlighted that this decline is not just a technical issue but a societal one, as misinformation campaigns increasingly use sophisticated AI to create convincing but false narratives. He urged media outlets to adopt multiโlayered verification, combining automated detection with human editorial review and sourceโtracing techniques.
Looking ahead, Pangram is investing in a new detection architecture that uses multimodal cuesโsuch as metadata, usage patterns, and crossโmodel comparisonsโto improve reliability. Spero said the company plans to release a beta API next quarter that will allow publishers to flag suspect content in real time. He also called on regulators to consider standards for AIโgenerated content, arguing that without clear guidelines, the risk of deepโfake text will only grow. The conversation around AI detection is gaining urgency as more platforms and governments grapple with the ethical and practical implications of machineโwritten prose.
Read Full Story at TechCrunch โ


