LLMs could write like humans but post-training guardrails make their text detectable

Bradley Emi, CTO of AI text detector Pangram, argues that post-training guardrails force large language models to produce detectable text despite their…

By Vane August 20, 2026 1 min read
LLMs could write like humans but post-training guardrails make their text detectable

Bradley Emi, CTO of AI text detector Pangram, argues that post-training guardrails force large language models to produce detectable text despite their theoretical capacity for human-like diversity. Systems such as ChatGPT, Claude, or Gemini adopt behavioural rules to avoid dangerous outputs or censor specific political statements. This process sharply narrows their expressive range, creating an effect known as mode collapse. Consequently, Pangram’s detection tools flag these constrained outputs but do not identify so-called base models, which retain greater variety before safety adjustments. The same applies to narrowly specialised fine-tunes trained only on specific authors or subreddits, as well as broken outputs like incoherent text. This limitation currently applies only to non-watermarked AI text, though watermarks will likely remain effective even with base model variety.

The practical implication is that safety mechanisms designed to prevent harm inadvertently create a digital fingerprint for AI-generated content. Detectors struggle to distinguish between a human writer and a model that has been heavily restricted by corporate safety policies. This creates a gap where standard safety training makes text easier to identify than raw model outputs.

* Base models write with more variety than safety-aligned versions
* Detection fails on narrowly specialised fine-tunes or incoherent text
* Watermarks remain effective regardless of model variety

Scroll to Top