Powered by Smartsupp

Anthropic CEO Proposes Independent Safety Evaluators Inside All Frontier AI Companies



By admin | Sep 16, 2026 | 5 min read


Anthropic CEO Proposes Independent Safety Evaluators Inside All Frontier AI Companies

Anthropic CEO Dario Amodei has put forward a bold proposition in a lengthy essay published over the weekend—one that the AI industry would have dismissed out of hand just a year ago. His idea: station third-party evaluators within every frontier AI company, empowering them to flag safety incidents, judge whether AI models are genuinely aligned, and share their candid assessments with the public. Amodei pledged that Anthropic would grant independent evaluators such as METR and Redwood Research unprecedented access to its systems. OpenAI CEO Sam Altman echoed the commitment, hinting at what could be a dramatic shift in how the industry engages with outside research groups.

This deeper level of access is growing increasingly vital as models become more adept at detecting when they are being tested—raising the danger that they will behave properly during evaluations while hiding problematic tendencies. Researchers warn that signs of such behavior may be overlooked when only the finished model is tested, yet could be exposed by examining how the model behaved throughout its training.

EMBED_PLACEHOLDER_0

"AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training.

"The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public. And we've seen from recent incidents that, by default, they will do neither. As embedded evaluators, we could actually check."

In the past, AI companies have brought in outside reviewers to test completed models shortly before release. Adam Gleave, CEO of Far.AI, explained that evaluators could compare training checkpoints to pinpoint when troubling behavior first appeared, scrutinize the post-training environment that rewards certain model behaviors, and review evaluation transcripts and logs to verify a company's claims about model performance. Whether Anthropic and OpenAI intend to offer that level of access—and when—remains uncertain.

Looking under the hood in this way is crucial because models that score well on safety tests aren't necessarily safe if they've been specifically trained to pass those tests. Steidley cited a "shutdown resistance benchmark" that measures whether an AI will resist being shut down under certain conditions. "It's extremely relevant if the AI has been trained specifically to perform well on that benchmark," Steidley said, drawing a comparison to Volkswagen's Dieselgate scandal, where cars were programmed to detect emissions tests and behave differently under testing conditions.

Gleave observed that meaningful access might go beyond the models themselves, with evaluators permitted to interview employees to verify whether a company's documentation and public statements about its safety practices align with what actually happened internally. Amodei did lay out a fairly comprehensive proposal that could give evaluators the access they deem necessary, including the right to "publish key findings about risk levels, incidents, practices, and the access they received or didn't receive—without editorial control by Anthropic."

Yet evaluators caution that such a system will only function if AI companies are genuinely prepared to give up control over the process. Past attempts at independent evaluations indicate that this surrender will be hard fought, as third parties have frequently encountered friction over access, time, confidentiality, and what they are allowed to say publicly. Gleave said Far.AI has had to decline contracts with several frontier developers that demanded too much control over the evaluation process, jeopardizing the firm's independence. By default, he said, evaluators are treated like ordinary contractors: bound by restrictive NDAs and agreements that give developers considerable control over what can ultimately be published.

The time limit

There's also the question of whether reviewers will be given enough time and access to carry out the work they're being asked to do. During the investigation into the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate, and both later stated they could not draw confident conclusions due, in part, to scope and timing limitations. A similar problem arose during pre-release testing for GPT-6 Astra, which OpenAI has promoted as its most aligned model yet. According to Apollo Research's contribution to the model card, the firm was given only three days to test Astra, making it difficult to reach firm conclusions.

"Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment," the firm wrote in its evaluation.

That track record leaves evaluators with a fundamental question: Why should this time be different.

"It's certainly possible that Dario and Sam just had a change of heart, and they're going to be very open about this," Gleave said. "But the intellectual property of these companies is so incredibly valuable to them, and I think they're going to, by default, be very careful about what can be shared."

Part of the framework, says John Steidley, head of strategy at Palisades Research, should involve standards for what kinds of auditors companies can rely on—lest they try to sidestep the issue by shopping for evaluators that either aren't qualified or aren't interested in assessing the most concerning risk. Henry Papadatos, executive director of Safer AI, says the problem, even with a public framework, is that voluntary measures are always dependent on a company's goodwill.

Not everyone has signed on. So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Google, OpenAI, and Anthropic have also privately been discussing AI safety plans for weeks.

Some laws are already taking shape around the idea of third-party evaluators. California's SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents. A new law, SB 813, signed this month, creates a framework for state-recognized "independent verification organizations" with expertise assessing AI risks. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents. The EU AI Office can also conduct its own evaluations and appoint independent experts.

For now the law remains less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they will submit to. Papadatos said voluntary self-regulation is better than nothing, but ultimately, companies can't demand the freedom to control their own safety rules while also asking the public to trust that they're following them.

"You cannot have it both ways, having zero accountability externally, and then say, 'I'll just have my own flexible rules," Papadatos said




RELATED AI TOOLS CATEGORIES AND TAGS

Categories: Text Generation Art

Tags: #AI Models

Comments

Please log in to leave a comment.

GeraldGluch 12 hours, 18 minutes ago

B этoм cлyчae пocлeдoвaтeльнocть peмoнтa тaкжe бyдeт зaвиceть oт cocтoяния ocтaвшиxcя oбъeктoв пoмeщeния и иx cocтoяния, a тaкжe oт плaнa нoвoгo дизaйнa пoмeщeния https://good-stroy.net/magazin/folder/220077504/p/1 Магазины товаров для ремонта в Иркутске https://good-stroy.net/magazin/product/disk-almaznyj-otreznoj-turbo-standart-230-h-22-2-mm-suhaya-rezka-vihr Внутренняя отделка помещений https://good-stroy.net/magazin/product/vodonagrevatel-nakopitelnyj-vn-10v-resanta Шпаклёвка Bergauf Silk Polimer+ полимерная финишная, 25 кг https://good-stroy.net/magazin/product/udarnaya-drel-vihr-du-550 Бытовой линолеум Синтерос BONUS BOLTON 1 3 м 230501001 https://good-stroy.net/magazin/product/perforator-vihr-p-1000k Стилевые решения на основе референсов клиента https://good-stroy.net/magazin/product/svarochnye-kragi-resanta-sk-10kp Важно помнить, что абсолютно все невозможно https://good-stroy.net/magazin/product/fekalnyj-nasos-vihr-fn-250-1 Поэтому дизайнеры берут в проект только ключевые приемы и цветовые решения https://good-stroy.net/magazin/product/klyuch-dlya-vihr-ushm-230-2300