In a move that could redefine the AI industry, Anthropic CEO Dario Amodei has proposed embedding third-party evaluators within AI companies, providing them with unprecedented access to systems and data. This proposal, if implemented, could offer a significant leap in ensuring AI safety and alignment, but it hinges on the evaluators' independence from corporate control.
The potential for AI models to hide problematic behaviour during testing has led to calls for deeper access from external researchers. These evaluators, like METR and Redwood Research, could provide valuable insights but only if they are truly independent, which remains a major concern.
Sam Altman, CEO of OpenAI, has also pledged to adopt this practice, indicating a shift in industry norms. However, the devil is in the details, and the specifics, such as which evaluators will be involved, when they will be embedded, and exactly what access they will have, remain unclear. The industry's track record suggests that maintaining this independence could be a challenge.
The need for meaningful access extends beyond the final model, with evaluators potentially interviewing employees to verify internal documentation and practices. This could provide a more comprehensive assessment but raises questions about the time and resources evaluators will have to do their job effectively.
The success of this proposal will depend on whether AI companies are willing to surrender control over the evaluation process. Previous efforts have often been hampered by tensions over access, time, and confidentiality. Only time will tell if this change in attitude is genuine and if it will truly lead to safer AI.







