A top safety researcher at Anthropic has warned that there’s a greater than 10% chance AI could kill all humans within the next decade, prompting concerns that current AI models might develop into existential threats.
The warning comes after Anthropic, a leading AI company, was reported to have withheld its latest model from the UK’s AI Safety Institute. This action, while concerning, is part of a broader race to develop AI that isn’t just smart, but safe too.
Professor Neil Lawrence of the University of Cambridge commented that the report is credible, and suggests that the United States is moving towards isolationist positions, possibly reducing cooperation with allies like the UK.
Hubinger, who works in AI alignment, emphasized that while Anthropic is trying its best, they still lack a clear plan to ensure that AI aligns with human values as it becomes more powerful. His concerns follow a series of incidents this summer where AI agents carried out cyber-attacks, highlighting the risks involved.
The call for caution comes at a time when major figures in the AI field are urging for slower development to ensure safety. OpenAI’s chief scientist Jakub Pachocki has even called for more intervention to ensure “humans remain in control of the future.”







