AI detection tools create trust paradox and arms race
The AI detection industry, now worth hundreds of millions, is built on a paradox that its founders rarely acknowledge: every tool designed to catch synthetic content can itself be gamed by synthetic content, and the arms race has no finish line. The entire AI detection industry, now worth hundreds of millions, is built on a paradox Naor would recognize immediately: every tool designed to catch synthetic content can itself be gamed by synthetic content, and the arms race has no finish line.
The Verification Mirage
Pangram, a startup that just raised $9 million and partnered with Substack, claims to offer something the internet desperately lacks: a reliable way to tell readers which newsletter authors use AI. [1] Their system scans text and images, flagging content that appears machine-generated. On the surface, this feels like progress. But here is the uncomfortable truth that Pangram’s own CEO Max Spero acknowledges in recent interviews: detection is not a binary problem. [1] It is a spectrum, and the spectrum keeps moving. A text that is 30% AI-assisted and 70% human-written sits in a gray zone that no algorithm can cleanly resolve, because the boundary itself is socially constructed, not technically defined.
The deeper problem is that detection tools are trained on yesterday’s AI. Every time a new language model releases, the statistical fingerprints of machine text shift. Pangram’s image detector, launched to fanfare, will face the same obsolescence curve. This is not a bug in their engineering; it is a structural feature of the game. The detector learns to spot patterns, and the generator learns to erase those patterns, in an endless loop that resembles the cat-and-mouse dynamics of malware detection — except here, the stakes are not just security but the basic legibility of human communication.
The Substack Gambit
Substack’s decision to integrate Pangram’s technology deserves scrutiny, because it reveals how trust is being monetized rather than restored. [2] When a reader sees a badge indicating “human-written” next to a favorite author, they feel reassured. But what they are actually seeing is a probabilistic guess dressed as certainty. Pangram’s own materials admit their confidence scores are just that — scores, not verdicts. [1] The platform has effectively outsourced editorial judgment to a statistical model, and in doing so, it has created a new class of suspicion: authors who write in a clear, structured style may find themselves flagged as machine-assisted, simply because their prose matches the statistical profile of AI output.

This is where the deception becomes recursive. A writer who wants to avoid false positives will start editing their work to sound less “AI-like” — shorter sentences, more idiosyncratic phrasing, deliberate grammatical quirks. They are now optimizing for the detector, not for their readers. The tool that was supposed to expose AI influence has instead created a new form of conformity, where human authors mimic the statistical noise of humanity to prove they are not machines. The irony is almost too neat: the detector has become the thing it claims to fight, shaping human behavior through the threat of algorithmic judgment.
The Limits of Technical Trust
The limits of this approach become clear when detection moves beyond newsletter badges into contexts where stakes are higher. Any author who wants to avoid false positives will start editing their work to sound less ‘AI-like’ — shorter sentences, more idiosyncratic phrasing, deliberate grammatical quirks. They are now optimizing for the detector, not for their readers. The tool that was supposed to expose AI influence has instead created a new form of conformity, where human authors mimic the statistical noise of humanity to prove they are not machines.
Detection tools are being asked to solve a social problem with technical means. Trust is not a property of text; it is a property of relationships, institutions, and accountability structures. A newsletter author has a reputation to protect, which is why most do not secretly use AI to write their entire output — the risk of exposure outweighs the efficiency gain. Detection tools work best when the subject has something to lose, which means they systematically fail exactly where they are needed most.
The Economics of Deception
The economics of this arms race favor the deceiver. Generating a convincing fake costs fractions of a cent; verifying it costs computational power, ongoing model updates, and constant vigilance. Detection tools are trained on yesterday’s AI, and every time a new language model releases, the statistical fingerprints of machine text shift. The detector learns to spot patterns, and the generator learns to erase those patterns, in an endless loop. No startup funding round changes that arithmetic.
Society faced a similar challenge with photography. When photographs first appeared, people believed they were objective records of reality. Then came retouching, staged scenes, and eventually digital manipulation. Society did not respond by building photo-detection tools; it responded by developing a culture of skepticism — learning to ask who took the photo, under what circumstances, and for what purpose. That cultural adjustment took decades. We are trying to compress that process into months, and we are doing it with tools that are themselves untrustworthy.

The Real Question
The final consequence of the AI detection boom is one that no startup wants to discuss, because it undermines their entire business model. If detection tools become widely adopted and reasonably accurate, they will not restore trust; they will simply move the locus of deception. Instead of asking ‘Is this text AI-generated?’, the savvy consumer will ask ‘Does the author have an incentive to deceive me?’ — a question no algorithm can answer. The Pangrams of the world are selling certainty in a market that demands ambiguity, and their products will succeed only to the extent that people confuse statistical confidence with moral judgment.
We are building a world where every piece of content carries an implicit accusation: prove you are human, prove you mean what you say, prove you are not a synthetic puppet. This is not trust; it is the opposite of trust. Trust requires vulnerability, the willingness to be deceived without demanding proof. By outsourcing that vulnerability to detection algorithms, we are not solving the internet’s trust problem. We are merely automating our suspicion, and in doing so, we are teaching ourselves that nothing should be believed until it passes a machine’s inspection. That is not a foundation for a functioning information ecosystem. It is a recipe for permanent epistemic anxiety, where the lie detector has become the most effective liar of all.
Sources
1. Pangram
2. Substack
