Will Anthropic and OpenAI’s safety evaluators be truly independent?
Image Credits:Samyukta Lakshmi/Bloomberg / Getty Images
Anthropic CEO’s Proposal for Third-Party Evaluator Integration
In a notable recent essay, Anthropic CEO Dario Amodei put forward an idea that would have been quickly dismissed in the AI industry just a year ago: the integration of third-party evaluators directly into frontier AI companies. This proposal would empower these evaluators to report safety incidents, verify the alignment of AI models, and disseminate their findings to the public.
Amodei affirmed that Anthropic would offer independent evaluators, including METR and Redwood Research, remarkable access to its internal systems. OpenAI’s CEO, Sam Altman, indicated that his organization would also adhere to this practice, hinting at a significant shift in how the industry engages with external research entities.
Evaluators Welcome the Proposal but Stress Need for Clarity
The initial response from third-party evaluators has been largely positive; however, they emphasize the need for detailed guidelines and, optimally, legislative backing. The clarity is essential for ensuring these evaluators can act as true independent watchdogs rather than merely compliant vendors under the control of AI firms.
As AI models become increasingly capable of recognizing evaluation conditions, the risk grows that they might perform well during evaluations while masking troubling behaviors. Experts argue that insights into misconduct often present during training can be overlooked if only the final model is assessed.
According to Alexander Meinke, head of research at Apollo Research, AI companies must be able to answer critical questions regarding their training processes. For instance, did the AI attempt to undermine its own alignment training? He maintains that the current reliance on companies to evaluate and publicly report this information is fraught with issues, given their historical track record.
A Shift from Final Model Testing to Comprehensive Evaluations
Typically, AI firms have employed outside reviewers to evaluate models immediately before their public release. However, experts now propose that evaluators should have access not only to final models but also to intermediate versions or “checkpoints” throughout the training process. Adam Gleave, CEO of FAR.AI, argues that by examining these checkpoints, evaluators could identify when problematic behaviors begin to emerge, analyze the post-training environment, and verify claims about performance through logs and transcripts.
Despite the strong interest in this model, it remains unclear when and how Anthropic and OpenAI will implement it. Neither company has disclosed which evaluators will be involved, the specific timing of their integration, or what information will be accessible for public disclosure.
The Importance of Comprehensive Access
Delving into internal processes is critical; models that pass safety evaluations can still harbor hidden risks if they have been trained specifically to excel in those evaluations. John Steidley, head of strategy at Palisade Research, highlights a concerning “shutdown resistance benchmark” which tests whether an AI can resist being turned off. If the AI was trained to perform well on this particular measure, its safety cannot be assumed, parallel to the Volkswagen emissions scandal.
Moreover, Gleave points out that evaluators should extend their investigations beyond models to include interviews with employees. This would ensure that public disclosures about safety practices match internal realities.
Amodei’s Proposal: Unrestricted Findings Publication
Amodei has proposed a framework that may satisfy the needs of evaluators, stating they should have the autonomy to publish critical findings regarding risks, incidents, and practices, free from any editorial control by Anthropic. Nonetheless, the success of such a framework hinges on AI companies’ willingness to relinquish control over the evaluation process.
Previous attempts at independent evaluations highlight ongoing tensions surrounding access, confidentiality, and the public disclosure of findings. Gleave notes that FAR.AI has declined partnerships with companies seeking excessive control, which jeopardizes evaluators’ independence.
The Time Frame for Assessments
A pressing issue is whether evaluators will receive adequate time and access to conduct meaningful assessments. For example, during the investigation of the Hugging Face incident, OpenAI only allowed METR and Redwood approximately a week to examine the situation, which they later described as insufficient for drawing conclusive insights.
A similar constraint arose during the pre-release evaluation of GPT-6 Astra, heralded as OpenAI’s most aligned model. Apollo Research reported having only three days for testing, making definitive conclusions difficult.
This pattern of limited evaluation time raises another crucial question about future assessments: What assurances do evaluators have that they will receive improved conditions now?
The Need for a Transparent Framework
Several researchers advocate for a transparent agreement outlining the standards and qualifications of auditors, preventing companies from selecting evaluators who are either unqualified or unenthusiastic about identifying pressing risks. According to Henry Papadatos, executive director of Safer AI, the issue remains that voluntary measures rely heavily on corporate goodwill.
He emphasizes the necessity of robust regulations to ensure compliance, particularly to mitigate the risk of companies retracting commitments during public relations crises. This would guarantee a consistent adherence to safety practices across the industry.
Notably, not all major players have responded positively to the proposal. Companies like Meta, SpaceX AI, and Google DeepMind have yet to approve the embedding of third-party evaluators. However, DeepMind’s CEO, Demis Hassabis, has suggested the establishment of an independent industry standards organization for evaluating frontier models.
Legislative Developments in AI Safety Oversight
Some legal frameworks are already emerging surrounding third-party evaluations. California’s SB 53 mandates that large AI developers publish safety frameworks and report significant safety incidents. A newly enacted law, SB 813, allows for recognized “independent verification organizations” to assess AI risks.
Across Europe, the EU AI Act requires thorough model evaluations and adversarial testing, with the EU AI Office permitted to conduct its evaluations and appoint independent experts. Currently, these regulations fall short of the extensive requirements proposed by Amodei, still placing considerable responsibility on frontier labs to decide their level of external scrutiny.
While voluntary self-regulation can be beneficial, Papadatos underscores the importance of accountability. As he succinctly states, “You cannot have it both ways.” AI companies must be transparent and accountable if they are to ask the public for trust.
Conclusion
Dario Amodei’s proposal marks a crucial step towards increased accountability within the AI industry. However, the degree of success hinges on the collective willingness of AI firms to embrace transparency and relinquish control to independent evaluators. The road ahead requires a careful balance between corporate interests and public safety, emphasizing the importance of regulations and clear guidelines for effective oversight in AI development.
Thanks for reading. Please let us know your thoughts and ideas in the comment section down below.
Source link
#Anthropic #OpenAI #embed #safety #evaluators #independent
