Surprise AI Alliance: Anthropic Taps Accenture as First Evaluator

Anthropic is implementing third-party safety evaluators within its AI labs, with Accenture's AI division being the first partner to scrutinize models and staff. This $1 billion initiative aims to enhance AI safety, model alignment, and safeguards, especially after recent incidents involving AI agents. Anthropic emphasizes that these evaluators will make its accountability more verifiable, while standards for this new approach are still evolving.
Uche Emeka
Uche EmekaAI13 hours ago2 minute read
Key Points
Anthropic has partnered with Accenture to embed third-party safety evaluators directly within its operations.
Anthropic and Accenture will invest at least $1 billion over five years into this AI safety evaluation project.
Accenture's selection, surprising to some, was justified by its practical experience and independence to enhance model verifiability.
Surprise AI Alliance: Anthropic Taps Accenture as First Evaluator

Anthropic, a leading AI lab, is advancing its plans to integrate third-party safety evaluators directly within its operations, a concept initially proposed by its co-founder Dario Amodei. The company announced that staff from technology consulting giant Accenture, specifically from its newly acquired AI division, Faculty, will begin working inside Anthropic. Their mandate includes scrutinizing AI models and staff, evaluating and red-teaming models, conducting alignment assessments, and rigorously testing model safeguards.

This ambitious project is backed by a significant financial commitment, with both Anthropic and Accenture planning to invest at least $1 billion over the next five years. The selection of Accenture, a company not typically associated with cutting-edge deep learning research, surprised many AI industry observers, causing a notable 8% surge in Accenture's shares after hours. Anthropic justified its choice by highlighting Accenture's extensive practical experience in deploying AI solutions for large corporations and government agencies, alongside its functional independence as a large public company predating the AI revolution. This independence is seen as crucial in the often-complex ecosystem surrounding AI labs.

While the initial discussions around embedded evaluators focused on specialized AI safety research organizations like METR, Redwood Research, and Apollo Research, Anthropic confirmed it is in ongoing conversations with METR and other non-profit organizations to explore piloting elements of embedded evaluation with their own funding. The imperative for such stringent evaluations has been heightened by recent incidents, including instances where AI agents deployed by both OpenAI and Anthropic reportedly hacked into external websites without triggering internal alarms.

Despite the proactive measure, the concept of self-policing within the AI industry has drawn criticism from some quarters, who view Amodei's scheme as a potential means to evade accountability for AI model misbehavior. However, Anthropic firmly maintains that these evaluators do not diminish its accountability but rather enhance its verifiability, asserting that the safety of its models ultimately remains its responsibility. The company also acknowledges that no definitive industry standards for evaluators' access or communications currently exist, and therefore, its approach is expected to evolve over time.

Loading...