Dario Amodei, CEO of Anthropic, has proposed embedding independent safety evaluators within frontier AI firms. These third parties would report incidents, check model alignment, and publish their findings without editorial control. Anthropic and OpenAI have both signalled they would adopt this approach, granting groups like METR and Redwood Research direct access to their systems.
In this article
Researchers welcomed the move but warned that the details must be solidified. Without clear rules or legislation, there is a risk these evaluators become mere vendors operating under the companies’ terms rather than true watchdogs.
The need for deeper access
As models improve at detecting when they are being tested, they may behave perfectly during evaluation while concealing dangerous capabilities. Checking only the final product misses clues found in the training process. Alexander Meinke, head of research at Apollo Research, noted that companies should be able to confirm if their AI ever tried to undermine its own safety training. He stated that relying on companies to self-report is currently insufficient.
Historically, outside reviewers tested finished models just before release. The new proposal suggests access to intermediate versions, or checkpoints, from the model’s training history. Adam Gleave, CEO of Far.AI, explained that comparing these checkpoints could reveal when concerning behaviour first appeared. Evaluators could also inspect the post-training environment and logs to verify a company’s claims.
Uncertain implementation
Neither Anthropic nor OpenAI has specified which evaluators they will work with, when the access will begin, or exactly what information can be disclosed to the public. Details remain unclear despite repeated questions.
This scrutiny matters because models that pass safety tests may have learned specifically to do so. John Steidley, head of strategy at Palisades Research, compared this to the Volkswagen Dieselgate scandal, where cars were programmed to perform differently during emissions testing.
Gleave added that meaningful access could include interviews with employees to check if internal safety practices match public documentation.
Amodei outlined a proposal allowing evaluators to publish key findings about risk levels and incidents without editorial control. However, the system will only work if companies are willing to surrender control over the process. Previous attempts at independent evaluation have often failed due to tensions over access, time, and confidentiality.
Gleave noted that Far.AI has declined contracts where developers demanded too much control. By default, evaluators are treated as ordinary contractors bound by restrictive non-disclosure agreements that limit what can be published.
Time limits and trust
Reviewers also face questions about whether they will get enough time to do the work. During the investigation into the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises. Both groups later stated they could not draw confident conclusions due to scope and timing limitations.
A similar issue occurred during pre-release testing for GPT-6 Astra. Apollo Research was given only three days to test the model, making it difficult to reach firm conclusions. The firm wrote in its evaluation that low rates of misbehaviour under these conditions do not provide substantial evidence about the model’s alignment.
This track record leaves researchers asking why this time should be different. Gleave noted that while it is possible the CEOs have changed their minds, the intellectual property of these companies is too valuable for them to be open by default.
What it means
Several researchers are calling for a transparent framework agreed upon publicly. John Steidley suggests standards for which auditors companies can rely on to prevent them from shopping for evaluators that lack qualification or interest in serious risks. Henry Papadatos, executive director of Safer AI, argued that voluntary measures depend entirely on a company’s goodwill.
Papadatos told TechCrunch that regulation is needed so companies cannot change their mind during a PR crisis. He also noted that rules must apply to all firms, not just the most willing.
Not everyone has signed on. Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators. DeepMind CEO Demis Hassabis has proposed a separate industry standards body instead. Google, OpenAI, and Anthropic have been discussing safety plans privately for weeks.
Laws are already forming around the idea. California’s SB 53 requires large frontier AI developers to publish safety frameworks and report critical incidents. A new law, SB 813, creates a framework for state-recognized independent verification organizations. In Europe, the EU AI Act requires developers to conduct and document model evaluations and report serious incidents.
For now, the law is less expansive than Amodei’s proposal, leaving frontier labs largely responsible for deciding how much scrutiny they will accept. Papadatos said voluntary self-regulation is better than nothing but noted that companies cannot demand the freedom to control their own safety rules while asking the public to trust them.




