A senior Anthropic safety researcher has estimated that artificial intelligence carries a greater than 10% chance of wiping out humanity within the next decade, hours after another researcher resigned from the Claude maker over concerns about the industry’s push toward increasingly powerful systems.
CNBC reported that Evan Hubinger, Anthropic’s Alignment Science Lead, said there was a greater than 10% chance AI could “kill all humans” within the next decade. His warning followed the resignation of Anthropic researcher Jacob Coxon, who accused leading AI labs of racing toward systems they may not know how to control.
The BBC similarly reported that Hubinger said the risk from models that currently exist remains “low”, but that he was worried future systems could become capable of improving themselves and move into far more dangerous territory.
Jacob Coxon quits over the race to superintelligence
Coxon, 27, said he had spent the past three years doing pretraining research at both OpenAI and Anthropic.
According to CNBC, Coxon said neither company is acting responsibly and accused them of racing toward self-improving superintelligence while “gambling with our lives.”
His concern centers on systems capable of recursively improving their own capabilities. Coxon argued that future superhuman systems could become powerful enough to hack computer systems, accelerate scientific progress and acquire meaningful access to resources.
The BBC reported that Coxon believes the people building advanced AI genuinely fear it could threaten humanity by the end of the decade.
That makes his resignation notable because the criticism is coming from inside one of the companies that has positioned AI safety and alignment as central to its identity.
Anthropic says it still lacks a superintelligence alignment solution
Hubinger largely agreed with Coxon’s description of the underlying risk.
Hubinger said Anthropic was trying its best but did not yet have a plan to solve alignment for superintelligence and was not clearly on track to develop one.
The BBC also highlighted Hubinger’s concern about recursive self-improvement, where increasingly capable systems could potentially contribute to improving successor systems and accelerate development beyond existing safety approaches.
His comments do not mean Anthropic believes today’s Claude models have a 10% chance of causing human extinction. Hubinger explicitly distinguished the relatively low risk he assigns to current systems from his concern about where frontier development could lead.
AI labs face a governance problem inside the engineering race
For AI builders, the dispute exposes a tension deeper than another argument about whether artificial intelligence is dangerous.
Alignment teams are trying to understand and constrain increasingly capable systems while the companies employing them are simultaneously competing on model performance, developer adoption and commercial scale.
That creates an engineering governance problem: what happens when capability development moves faster than the techniques designed to understand or control it?
Coxon’s resignation and Hubinger’s public estimate do not prove catastrophic outcomes are inevitable. The 10% figure is Hubinger’s personal risk assessment, not a measured probability.
But when researchers working directly on frontier-model alignment publicly assign a non-trivial probability to an existential outcome—and acknowledge that the technical solution is unfinished—the industry faces a difficult question about how much uncertainty is acceptable while continuing to scale.