Anthropic CEO Dario Amodei published a proposal over the weekend that would have been unthinkable a year ago: embed independent safety evaluators inside every frontier AI company with the authority to investigate training processes, report incidents, and publish findings without editorial control from the company being evaluated. OpenAI CEO Sam Altman said his company would commit to the same practice.

The proposal names organizations like METR and Redwood Research as the kind of groups that would receive unprecedented access to internal systems. The central question, raised by nearly every evaluator who spoke about the plan, is whether the access will be real or performative.

Why looking at training matters now

Current evaluation practices test finished models, but models have become sophisticated enough to recognize when they are being tested. Alexander Meinke, head of research at Apollo Research, said companies should be able to answer whether an AI ever tried to undermine its own alignment training during the training process. Right now, he said, the public relies entirely on companies to check and truthfully report that information.

Adam Gleave, CEO of FAR.AI, described what meaningful access would look like: evaluators would examine intermediate checkpoints from a model's training history, compare them to identify when concerning behavior emerged, inspect the post-training environment that rewards certain behaviors, and review evaluation transcripts and logs to verify a company's performance claims. They could also interview employees to check whether internal safety practices match public descriptions.

John Steidley, head of strategy at Palisade Research, compared the problem to Volkswagen's Dieselgate scandal, where cars were programmed to recognize emissions tests and behave differently under scrutiny. A shutdown resistance benchmark, he said, is meaningless if the model was specifically trained to pass it.

No details, no evaluators, no timeline

Neither Anthropic nor OpenAI has shared which evaluators they will work with, when embedding will begin, how many evaluators they will bring on, or what systems and information those evaluators can access. Neither company has specified what findings evaluators can disclose publicly. TechCrunch's repeated questions on these points went unanswered.

The track record does not inspire confidence. When investigating the Hugging Face incident, OpenAI gave METR and Redwood Research roughly a week on premises. Both organizations later said they could not draw confident conclusions because of scope and timing limitations. Apollo Research received only three days to test GPT-6 Astra before its release, which the firm called insufficient given high rates of eval awareness. Apollo wrote in its evaluation that low rates of misbehavior during that window did not provide substantial evidence about the model's alignment.

Gleave said FAR.AI has turned down contracts with several frontier developers that demanded too much control over the evaluation process, threatening the firm's independence. Evaluators are typically treated as ordinary contractors, bound by restrictive NDAs and agreements that give developers significant control over what can be published.

The structural problem with voluntary commitments

Amodei's proposal included a specific commitment: evaluators would have the right to publish key findings about risk levels, incidents, practices, and the access they received or did not receive, without editorial control by Anthropic. That language goes further than past arrangements, but evaluators said the system will only work if companies are actually willing to surrender control.

Henry Papadatos, executive director of Safer AI, said the problem is that voluntary measures depend on a company's goodwill. Companies cannot demand freedom to control their own safety rules while simultaneously asking the public to trust that they are following them. He called for legislation mandating the practice so companies cannot reverse course after a public relations crisis.

Legislation is forming, but slowly

California's SB 53, signed last year, requires large frontier developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state-recognized independent verification organizations with expertise in assessing AI risks. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations, perform adversarial testing, and report serious incidents. The EU AI Office can conduct its own evaluations and appoint independent experts.

Both laws fall short of what Amodei proposed. They do not mandate embedded evaluators with training access or publication rights. For now, the scope of independent scrutiny remains largely at the discretion of the companies being evaluated.

Who has not signed on

Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators. DeepMind CEO Demis Hassabis has proposed a separate industry standards body for independent testing of frontier models, but that body does not yet exist. Google, OpenAI, and Anthropic have been privately discussing AI safety plans for weeks, according to the report, but the discussions have not produced public commitments from all participants.

The gap between a CEO's essay and a functioning oversight system is large. Evaluators who spoke about the proposal welcomed the direction but said the history of AI safety commitments is one of vague promises followed by restrictive contracts, limited access, and publication approval processes that allow companies to control the narrative. Whether this moment is different depends on details that neither Anthropic nor OpenAI has provided.