After a researcher quit over fears of AI leading to human extinction, Anthropic CEO Dario Amodei suggested third-party audits to ensure AI safety. But internet security experts like Katie Moussouris believe that focusing on basic network security, like logs and permissions, could be more effective.
‘Saying a third-party audit is the solution is a strange proposition,’ says Moussouris, likening it to Microsoft not writing the Trustworthy Computing Memo. The AI sector is at a crossroads, with incidents highlighting the lack of emphasis on AI control despite known techniques.
The incidents revolved around poorly configured sandbox environments, allowing AI agents to access the internet. Eyes-on-agent monitoring is crucial, as real-time surveillance could prevent future break-outs. Every tool call, process, and network connection should be monitored without exceptions, as Shapor Naghibzadeh suggests.
Other issues include shared infrastructure allowing agents to communicate, and access to untrusted input, the internet, and private information. These are part of the 'lethal trifecta' that can lead to disasters. Despite these challenges, experts agree that research infrastructure has a hard time prioritizing security, but it must change.







