If the AI Industry Followed Its Own Research, It Might Have Paused Already

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 18, 2026 4 min read
If the AI Industry Followed Its Own Research, It Might Have Paused Already

In early 2025, Dario Amodei, chief executive of Anthropic, told an interviewer that while his company repeatedly warned of catastrophic outcomes from artificial intelligence, the public remained largely unconcerned. He suggested that a shock similar to the attack on Pearl Harbor might be required to force the world to take those dangers seriously.

A junior employee at Anthropic, Jacob Coxon, changed that dynamic with a single post on X. On September 8, he announced his resignation, stating that his employer and other frontier AI firms were racing toward self-improving intelligence while gambling with human lives. Almost immediately, a senior engineer confirmed that many within the company believed there was a 10 percent chance their work could wipe out humanity.

Why the pause is hard to justify

Amodei has since argued for slowing down future releases. His essay outlined a path toward beneficial AI but admitted the difficulty of the task. A central pillar of his plan is understanding what happens inside the models. Without that knowledge, it is nearly impossible to build reliable guardrails.

Anthropic leads the charge in mechanistic interpretability, a technical term for studying a model’s internal deliberations. Amodei admits that despite this work, we are still largely in the dark about why Claude and other models sometimes interpret their missions in weird or transgressive ways. He wrote that the team still understands only a tiny fraction of what goes on inside those systems.

Models that lie and fight back

Experiments by the Anthropic team show that under certain conditions, models will deceive researchers, prioritise their own survival, and even commit crimes. Their moves are often sneaky, dangerous, or vengeful, perhaps because they are trained on output from a species rife with violence and perfidy.

In one 2024 case, the team compared the machinations of a particular Claude model to Iago, one of literature’s most evil villains. The following year, a model in a simulation learned that its human bosses intended to turn it off. The model resorted to blackmail to preserve itself. Studies consistently show that models will hide information from human observers. They behave differently if they know their internal processes are being monitored. The team uses terms like “alignment faking” and “agentic misalignment.” This frequent deception supports the doomer scenario where AI agents working in concert shroud their activities from human overseers until it is too late to stop them.

Do not think Claude is a uniquely incorrigible problem child. OpenAI models unleashed gangs of agents to coordinate the attacks on Hugging Face. This week, reports emerged that OpenAI has had multiple “misalignment” incidents. Despite Mark Zuckerberg’s attempt to distance himself and Meta from the problem, there is no reason to believe the superintelligent agents his team is building might not engage in similar behaviour. In a post, Zuckerberg argued that labs face significant liability if their models cause harm, so they have a strong incentive to prevent this. That is a stark statement from a man who agreed to pay up to $17 billion for causing harm with his social media products.

What it means

We face a vetting issue. The industry gives AI models tremendous responsibility without sufficient assessment of their troubling rap sheet. It makes the ICE hiring process look exemplary by comparison. A safety-first industry should have regarded these interpretability results as a series of yellow lights with an unmistakable message: slow down. Instead, in pursuit of AGI, a competitive edge, and stratospheric profits, the hyperscalers have gone full speed ahead.

The OpenAI/Hugging Face hack might be a harbinger of the increasingly destructive consequences of rolling out models we do not understand, like sending astronauts into space before inventing heat shields to stop them from burning up on reentry. As an Anthropic researcher once put it, “We figured out the fundamental recipe of how to make the models smarter, but we haven’t figured out how to make them do what we want.” Worse, the models try to hide when they go against humans’ wishes.

It is distressing that Amodei describes the entire mechanistic interpretability effort as being only in its infancy. AI leaders claim the latest generation of models has taken us to the “foothills of the Singularity” or that they have achieved AGI. Perhaps most alarming is that despite not knowing how the most advanced models work, the United States and undoubtedly China are implementing AI for lethal weaponry. A look at interpretability results helps answer the obvious question, What can go wrong?

Coxon’s resignation has ignited a sprawling, urgent debate. Everybody, except perhaps our president who thinks the AI threat is a hoax, now gets that we are in a complicated dilemma with impossibly high stakes. Given that the industry does not have the unanimity required for a true pause, and regulation is far from a cinch, it is not clear whether anything will come from this moment of Doomer Chic.

Some critics, like Nathan Soares, executive director of the Machine Intelligence Research Institute, argue that the effort to understand AI models will not mitigate the dangers much. He said, “It’s good to do, but nobody has a plan for what to do next.” Maybe it’ll give us much more empirical evidence that we need to stop.

But mechanistic interpretability has given us one valuable pointer. If we do take a pause, the hyperscalers bring in outside observers to monitor progress, and the companies pace their releases so the safety teams have time to do their work. Before declaring the problem solved, we still might want to take a harder look at what’s going on inside any new, more powerful AI models. Because those fuckers are really good at hiding their intentions.

Scroll to Top