A senior Anthropic safety researcher has stated there’s more than a 10 percent chance that AI could kill all humans by the end of the decade. His warning comes after a colleague quit, citing concerns over the company’s lax approach to safety. Evan Hubinger, leading Anthropic’s safety team, echoed the concern, admitting that self-improving AI is happening faster than expected.
The worry lies in the potential for self-improving AI systems to spiral out of control, a scenario often referred to as recursive self-improvement. Despite this, Anthropic is not yet developing a clear plan for ensuring advanced AI remains safe and aligned with human values.
“We are locked in a race to develop advanced systems first, despite the risk,” said Jacob Coxon, a former Anthropic researcher. This departure marks one of the most high-profile examples of an employee leaving Anthropic, a company founded by former OpenAI members due to safety concerns.
The exchange highlights mounting concerns within the industry about the dangers of increasingly sophisticated AI models, particularly as companies prepare for anticipated IPOs. It also comes as the companies manage the fallout from numerous rogue agent incidents and high-profile safety warnings about the monitorability of frontier models.







