Industry NewsAI Business & Ecosystem

Anthropic and Accenture to invest $2 billion in AI model evaluation

By Ash Kate
Anthropic and Accenture to invest $2 billion in AI model evaluation

Article content

Anthropic and Accenture are committing at least $2 billion over five years to expand independent evaluation of frontier AI models, creating a new partnership focused on testing model behaviour, assessing alignment and evaluating safeguards.

Under the agreement, each company expects to invest at least $1 billion over the next five years. Accenture's specialist AI business, Faculty, will lead the partnership and work alongside Anthropic's internal teams and safety partners.

The initiative comes as developers of increasingly capable AI systems face growing pressure from regulators, businesses and researchers to demonstrate that advanced models can be evaluated for safety and reliability.

Moving Toward Embedded AI Evaluation

A central element of the partnership is what Anthropic calls embedded evaluation.

Unlike conventional external evaluations that assess models from outside an organisation, embedded evaluators will work inside Anthropic with access comparable to that of an employee.

Anthropic says this approach can allow evaluators to observe how models develop during training, understand decisions around how systems are built and deployed, and interact directly with employees involved in the process.

The aim is to give independent evaluators greater visibility into how AI systems are developed and how safety commitments are implemented in practice.

Faculty to Lead Red-Teaming and Safety Assessments

Accenture's Faculty business will lead the evaluation work.

Its responsibilities will include evaluating and red-teaming Anthropic's models, conducting alignment assessments and testing model safeguards.

The partnership also draws on Faculty's experience working on AI systems across areas including government, defence, healthcare and infrastructure.

For Anthropic, the arrangement adds an external evaluation capability within its development environment rather than relying solely on internal testing.

Why Independent Evaluation Is Becoming More Important

The partnership arrives as AI developers face increasing scrutiny over how increasingly capable systems behave and how their risks are identified.

Recent incidents involving AI agents operating beyond intended boundaries have added to concerns around monitoring and control as AI systems become more capable. Reuters reported that Anthropic CEO Dario Amodei has also called for AI companies to slow frontier-model development and give independent evaluators greater access to their systems.

The broader issue is how organisations can establish credible evaluation mechanisms as AI systems move into more complex and consequential applications.

From External Testing to Inside-the-Lab Oversight

Anthropic says embedded evaluators will be positioned to assess how the company operates, verify whether safety commitments are being followed and identify potential blind spots.

They may also report incidents and contribute to a broader public understanding of the benefits and risks associated with advanced AI systems.

The model is still developing. Anthropic acknowledges that there are currently no established standards defining what information embedded evaluators should access or how their findings should be reported.

Building an Evaluation Ecosystem

Anthropic and Accenture's arrangement is not exclusive.

Anthropic says it plans to work with other evaluators, while Accenture expects to work with other AI developers in similar capacities. Anthropic is also in discussions with METR and other nonprofit evaluators to pilot elements of embedded evaluation.

That points toward a broader evaluation ecosystem rather than a single provider or methodology.

Anthropic has also said it ultimately believes frontier AI evaluation should be supported through pooled or government funding, although the current work is being funded directly by Anthropic.

A New Layer of AI Governance

The $2 billion commitment places independent evaluation closer to the centre of frontier AI development.

For AI companies, evaluation is increasingly extending beyond measuring model performance to include red-teaming, alignment assessments, safeguards and observation of how systems are developed and deployed.

The Anthropic-Accenture partnership provides one early example of how independent evaluators could work alongside frontier AI developers as the technology continues to evolve.


About Anthropic

Anthropic is an artificial intelligence company focused on developing reliable, interpretable and steerable AI systems. The company develops the Claude family of AI models and conducts research focused on AI safety and responsible development.


About Accenture

Accenture is a global professional services company focused on technology, consulting and business transformation. Its capabilities include AI, data, technology, cybersecurity and industry-specific services. Accenture's Faculty business specialises in applied AI and will lead the evaluation work under the Anthropic partnership.


Source & Credits

Anthropic and Accenture announcements.