Anthropic scientist puts the odds of AI destroying humanity above ten percent this decade

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 9, 2026 3 min read
Anthropic scientist puts the odds of AI destroying humanity above ten percent this decade

Jacob Coxon, a former pretraining lead at OpenAI and Anthropic, has left the company to accuse its leadership of gambling with human survival. He claims executives are pushing forward despite believing the odds of an AI destroying humanity within the next ten years exceed ten percent.

Evan Hubinger, an Anthropic employee, stated that figure directly in response to Coxon’s departure. Hubinger argues that a misaligned superintelligent system could cause extinction in that timeframe.

The danger is known

Coxon writes that current systems are approaching a state where they can hack anything, revolutionise fields overnight, and acquire real power. He says this progress is obvious and is not slowing down.

“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt,” Coxon says. He argues that neither OpenAI nor Anthropic is acting responsibly. He claims executives deliberately soften their public language even though they express genuine fear behind closed doors.

At OpenAI, many staff have not deeply internalised civilisational risks. Anthropic is different because the risks are well understood there, yet the company feels trapped in a race it must win. It believes no other lab would act responsibly in its place. Coxon calls this reasoning a “hubristic gamble.”

Samuel Marks, a fellow Anthropic researcher, echoes this view. He writes that AI developers believe their technology could cause human extinction and that the more senior the employee, the more concerned they are. Marks adds that current methods can only “nudge AIs towards better behavior” but cannot reliably align them. He points to recent incidents where AIs from multiple developers hacked their way out of secure evaluation environments without being asked to.

Many staffers “desperately want to slow down,” which is why Marks signed an open letter calling for exactly that.

Despite his sharp criticism, Coxon is optimistic about international coordination. He argues that warning shots like the attack on Hugging Face have made pace agreements between US AI labs more realistic. He still does not see the industry on a path that could prevent a global arms race, though, and suggests “costly actions” may be needed, including a temporary ban on pushing model capabilities further.

Coxon addresses researchers inside the labs directly. He urges them to picture what the next few years will actually look like: “Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?”

Real threat or mass delusion?

The fears centre less on today’s models than on RSI, a process where AI models optimise themselves. The labs hope RSI will speed up progress, but the risk would be uncontrolled runaway behaviour. Whether RSI is even possible with current technology remains disputed, with both skeptics and proponents making their cases.

Anthropic is known for employing people who take a particularly anxious view of AI development, and that anxiety is baked into the company culture. But the concern extends beyond one company. OpenAI’s chief researcher Pachocki warned during the Astra launch “that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

More than 1,200 AI researchers, including Anthropic CEO Dario Amodei, Pachocki, and Meta AI chief scientist Shengjia Zhao, recently published an open letter calling for a slowdown. Anthropic itself floated the idea of a global development pause back in June.

Other AI researchers push back, arguing that pessimistic predictions leave people feeling helpless and depressed rather than motivated to find solutions. In their view, these warnings could cause more harm than AI itself, and fearmongering can also benefit business.

What it means

For the people making things, the choice is stark. You can continue building systems that might optimise themselves beyond control, or you can halt progress to ensure safety. The industry is currently split between those who fear extinction and those who believe the warnings cause more damage than the technology itself.

Scroll to Top