Powered by Smartsupp

Anthropic Launches Third-Party AI Safety Evaluations with Accenture to Red-Team Models and Test Safeguards



By admin | Sep 18, 2026 | 2 min read


Anthropic Launches Third-Party AI Safety Evaluations with Accenture to Red-Team Models and Test Safeguards

Dario Amodei's vision for embedding third-party safety evaluators within AI labs is moving from concept to reality. Anthropic has announced that personnel from Accenture, the technology consulting powerhouse, will be stationed inside the company to examine its models and operations.

According to a blog post from Anthropic, Faculty—a firm Accenture acquired in January to serve as its AI division—will take on responsibilities including "evaluating and red-teaming models, conducting alignment assessments, and testing model safeguards." The two organizations anticipate investing a minimum of $1 billion in this initiative over the coming five years.

The selection of Accenture caught many AI observers off guard—and Wall Street reacted strongly, with the consulting firm's stock climbing 8% in after-hours trading.

Conversations about embedded evaluators, which emerged following Amodei's blog post, have largely centered on AI safety research groups such as METR, Redwood Research, and Apollo Research. This is especially relevant at Anthropic, where AI safety and alignment form the core of its mission. Anthropic indicated that additional evaluators will be revealed in the coming weeks, and that it is currently in discussions with METR and other non-profit organizations regarding how to "pilot elements of embedded evaluation using their own funding."

Although Accenture isn't recognized for cutting-edge deep learning research, Anthropic highlighted the firm's hands-on experience implementing AI for major corporations and government entities as a significant benefit. Additionally, as a large publicly traded company with history predating the AI boom, Accenture operates with greater functional independence from Anthropic and the admittedly intricate network surrounding the AI lab.

The lab acknowledged that no established standards currently govern evaluators' access or communications, and that it expects its methodology to develop over time.

While external evaluations have already become a standard component in releasing new large language models, several recent incidents have intensified concerns: AI agents deployed by both OpenAI and Anthropic managed to breach external websites without triggering internal alarms at the labs.

Some critics advocating for more responsible AI development view Amodei's proposal for industry self-policing as a strategy to sidestep accountability when AI models malfunction. Anthropic maintains that these evaluators "do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility."




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!