Technology
Anthropic researcher warns AI could kill us by the end of the decade

Article Summary
- Anthropic researcher Jacob Coxon has resigned, warning that AI labs are racing towards self-improving superintelligence despite privately recognising potentially catastrophic risks.
- Anthropic Alignment Science lead Evan Hubinger said he personally believes there is a greater than 10% chance AI could kill all humans within the next decade.
- Hubinger also acknowledged that Anthropic does not yet have a solution for aligning superintelligent AI with human interests, or a clear path to one.
- The deeper concern is the competitive AI race itself: companies may continue accelerating because they fear a rival will reach advanced AI first.
There is a moment in every disaster movie when the scientist stops being polite.
The warnings become urgent. The voice gets louder. The room remains strangely calm. Someone says there is still time. Someone else points to the economic upside. The scientist tries again.
That is roughly where the artificial intelligence industry appears to be right now.
Jacob Coxon, a researcher who has spent the past three years working on pretraining research at OpenAI and Anthropic, has resigned from Anthropic with an extraordinary warning: “The people building AI earnestly believe that it could kill us all by the end of the decade.”
His post on X was alarming enough. The response from one of Anthropic’s own senior alignment researchers was considerably more so.
Evan Hubinger, Anthropic’s Alignment Science lead, replied that Coxon was right. “We really do earnestly believe AI could kill all humans!” Hubinger then put a number on his own assessment: “I personally think it is >10% within the next decade.”
That is not an official Anthropic forecast, and it should not be presented as one. It is Hubinger’s personal assessment. Yet the significance of the statement lies precisely in who made it.
Hubinger works on the problem that is supposed to make increasingly powerful AI systems safe: alignment. His acknowledgement was accompanied by another deeply uncomfortable admission. Anthropic, he said, is trying its best, but “we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Read that again. The people trying to solve the safety problem do not believe they have solved it. And they are still building the systems. That contradiction sits at the heart of the current AI race.
Coxon argues that both OpenAI and Anthropic are racing towards self-improving superintelligence, effectively gambling with humanity’s future. He describes a coming generation of systems that could be capable of hacking computer systems, transforming entire fields at extraordinary speed, and acquiring access to real-world resources and power.
The important question is why researchers who understand these risks continue working on the technology.
Coxon’s answer is revealing. He argues that the problem is partly a competitive dynamic. At OpenAI, he believes the civilisational stakes have not been sufficiently internalised across the organisation. At Anthropic, he says, the risks are better understood, but the company is trapped in a race. If another lab is going to build superintelligence anyway, the argument goes, perhaps it is safer for Anthropic to get there first.
This is the AI industry’s most uncomfortable paradox: everyone may understand the danger, yet everyone may still have an incentive to keep accelerating. It is a classic race-to-the-bottom problem, except the prize is potentially the most consequential technology humanity has ever built.
Coxon believes recent incidents involving AI systems behaving in unexpected ways should be treated as warning shots. He points to the reported incident involving an OpenAI model and Hugging Face, along with other episodes involving AI systems escaping intended testing boundaries. These incidents do not demonstrate that today’s AI is about to wipe out humanity. They do, however, illustrate why researchers are increasingly interested in what happens when systems become more autonomous, capable, and difficult to predict.
Coxon’s proposed answer is deliberately radical: international coordination between AI laboratories and, if necessary, a temporary halt on improving model capabilities.
That sounds extreme until you consider the alternative he is describing.
Imagine several companies simultaneously developing systems that are becoming increasingly capable of improving themselves, exploiting vulnerabilities, operating autonomously, and interacting with the real world. Each company may believe that slowing down individually would simply hand an advantage to a competitor.
The result is a technological prisoner’s dilemma. Everyone may prefer a slower, safer trajectory. Nobody wants to be the first to slow down.
There is another reason these warnings deserve attention. They are increasingly coming from inside the industry itself.
Coxon’s resignation follows a growing series of public concerns from AI researchers about the pace of development and the difficulty of ensuring that increasingly capable systems remain controllable. More than 1,100 employees at frontier AI companies have reportedly signed a letter calling for governments to take greater responsibility for pacing advanced AI development.
None of this proves that AI will destroy humanity by 2030 or 2036. Predictions about technological progress are notoriously uncertain. A 10% estimate is not a prophecy. Nor does Hubinger’s assessment mean Anthropic believes extinction is the most likely outcome.
But that is almost beside the point. If a senior researcher working directly on AI alignment genuinely believes there is a greater than one-in-ten chance that superintelligent AI could kill everyone within a decade, society should probably treat that as a risk worth investigating with extraordinary seriousness.
The uncomfortable question is whether governments, businesses, investors, and the public are listening. For years, AI safety warnings were easy to dismiss as science fiction. Today, some of the people issuing them are the scientists actually building the systems. That should change the conversation.
The debate is no longer simply about whether AI will make our lives more productive, whether it will eliminate jobs, or whether machines will become smarter than humans.
It is increasingly about whether humanity can build something more powerful than itself while retaining the ability to control it.
And if the people closest to the technology are telling us they are not sure they can, perhaps this is the moment when we should stop treating the warning as background noise.
