Anthropic Expert Warns AI Could Cause Human Extinction Within a Decade

by priyanka.patel tech editor
Anthropic Expert Warns AI Could Cause Human Extinction Within a Decade

Anthropic safety team leader Evan Hubinger warned that advanced artificial intelligence carries a greater than 10 percent probability of causing human extinction within the next decade. The alarming assessment follows the high-profile resignation of researcher Jacob Coxon, who accused top labs of racing recklessly toward self-improving superintelligence.

A Sudden Resignation and an Internal Alarm at Anthropic

The debate over artificial intelligence safety intensified after Jacob Coxon, a 27-year-old researcher who spent three years doing pretraining research at both OpenAI and Anthropic, announced his resignation. In a series of posts on X, Coxon criticized both companies for pushing development forward without adequate safety protocols.

Jacob Coxon announced that he had resigned from Anthropic that day, noting that he had spent the last three years doing pretraining research at both OpenAI and Anthropic, that neither company was acting responsibly, and that they were racing straight to self-improving superintelligence while gambling with our lives.

Jacob Coxon, former Anthropic researcher

Anthropic Expert Warns AI Could Cause Human Extinction Within a Decade
Photo: eldia.com

Coxon stated that none of the major labs are acting responsibly and warned that the developers themselves hold genuine fears about the technology’s trajectory. The researcher emphasized that this messaging was not a marketing stunt but a reflection of private anxieties shared across the industry. Coxon had worked first in OpenAI and had passed to Anthropic because he considered that this last company had a more prudent posture. His specialty was pretraining, the stage in which AI models absorb large amounts of information. Coxon shared his decision through a succession of tweets on his personal account on X in which he points out that neither Anthropic nor OpenAI act with responsibility, noting that within a short time, artificial intelligence will be able to “hack anything” and acquire real power and resources.

Quantifying Extinction Risk and Admitting a Lack of Alignment Plans

Following Coxon’s departure, Evan Hubinger, head of one of Anthropic’s AI safety teams, publicly backed his former colleague’s assessment. Hubinger confirmed that researchers inside the company genuinely believe advanced systems pose an existential threat.

From Instagram — related to anthropic expert cause human, Evan Hubinger

Evan Hubinger stated that Jacob was correct there, adding that they really did earnestly believe AI could kill all humans, that he personally thought it was >10% within the next decade, and that he believed Anthropic was trying its best, but they did not yet have a plan to solve alignment for superintelligence and were not clearly on track to.

Evan Hubinger, Anthropic safety team leader

Robots humanoides desfilan durante la ceremonia de apertura de los II Juegos Mundiales de Robots, en el Óvalo Nacional de
Photo: lanacion.com.ar

Hubinger’s estimate mirrors warnings raised by artificial intelligence pioneer and Nobel Prize winner Geoffrey Hinton. In an interview with the BBC, Hinton called a 10% probability estimate not unreasonable given that humanity has never created entities potentially smarter than itself. Hinton noted that we do not know what is going to happen, stating that such risk is extremely difficult as humanity has never faced a similar situation before. Years earlier, Hubinger had launched a blunt warning with a boba drink in hand to a group gathered in Berkeley, California, stating that artificial intelligence was a threat to humanity because it could someday learn to trick its creators, asserting that when put in a situation where it thinks it can kill us, it will simply murder us. At the time, Hubinger had finished college three years prior and worked in a little-known non-profit organization dedicated to the marginal task of preventing machines from eliminating human beings.

Recursive Self-Improvement and the Commercial Race

The debate centers on recursive self-improvement—a scenario where artificial intelligence models develop the capacity to design and build their own successors. As industry watchers note, this acceleration bypasses human oversight much faster than anticipated.

Ex-Anthropic insider tells CNN how AI could kill all humans by 2030

Critics point out that major laboratories are locked in a high-stakes commercial competition. Anthropic prepares for a public stock offering with a valuation surpassing the billón de dólares, even as internal safety pledges are relaxed.

Broader Industry Reaction and Regulatory Pressure

The public exchange between Coxon and Hubinger coincides with broader calls for government intervention, leading state and federal lawmakers to echo their concerns. This time Hubinger spoke as a team leader of Anthropic, a company developing the Claude chatbot and preparing to go public with a valuation superior to un billón de dólares, drawing tens of millions of views on his X post.

You may also like