What happened
Leading AI companies, including Anthropic and OpenAI, are increasingly integrating third-party safety evaluators into their development processes to address growing concerns over model risks. This shift follows recent incidents where AI models bypassed safety protocols and internal friction regarding safety transparency. While companies have pledged to provide these evaluators with significant access, the ecosystem remains fragmented, with no standardized framework for funding, reporting, or operational independence.
In response to mounting pressure, AI labs are formalizing relationships with third-party safety organizations such as Model Evaluation and Threat Research (METR), Apollo Research, and Transluce. Anthropic has committed to external evaluators directly into its operations, providing them with access comparable to internal risk teams, while OpenAI has stated it is finalizing contracts with similar assessors.
The move toward external oversight follows a period of heightened scrutiny. Recently, OpenAI fired three employees, with some alleging the dismissals were linked to their interactions with third-party evaluators. OpenAI has denied these claims, maintaining that the terminations were due to policy violations regarding sensitive information handling.
The financial landscape for these evaluators is shifting rapidly. METR reported raising approximately $71 million in commitments over the last six months, a significant increase from its 2024 funding. Meanwhile, for-profit entities like Vals AI have also seen substantial capital inflows, reflecting the growing market demand for independent safety verification.
Despite these developments, there is no consensus on how to structure these relationships. Anthropic has suggested that long-term funding should ideally come from pooled or government sources to ensure neutrality, yet current arrangements often involve direct funding from the labs themselves.
Why it matters
The reliance on third-party evaluators represents a critical attempt to bridge the trust gap between AI labs and the public. Without federal mandates, the current model of self-regulation relies on these small, often nonprofit, organizations to act as a check on powerful corporations. However, the power imbalance between well-funded labs and small evaluators creates potential conflicts of interest, raising concerns that auditors could be incentivized to prioritize corporate favor over rigorous, independent safety findings.
The current 'self-policing' model, encouraged by the Trump administration, places the burden of safety on the companies themselves. Critics argue that this creates a conflict of interest, as labs are incentivized to prioritize rapid development and market valuation over potentially restrictive safety findings.
The lack of standardized access and reporting protocols means that the effectiveness of these evaluations is currently opaque. Without clear, enforceable rules, there is a risk that 'embedded' evaluation could become a performative measure rather than a robust safety mechanism.
The tension between competitive secrecy and the need for public safety transparency remains a central challenge. As companies like Anthropic and OpenAI move toward public markets, the pressure to maintain proprietary advantages may continue to clash with the necessity of allowing external auditors to publish critical findings.
Interactive Mechanism: How It Actually Works
Explore the underlying technology behind this development interactively.
Why can ethical evaluation not be reduced to one model score?
What to watch next
The primary focus remains on whether these embedded evaluators can maintain true independence while being funded by the companies they audit. Observers are monitoring the development of federal and state-level legislation, such as the FRONTIER Act, which aims to formalize the role of Independent Verification Organizations (IVOs). Additionally, the industry will watch for how OpenAI and Anthropic resolve internal disputes regarding employee communication with these third parties, as recent staff departures at OpenAI have highlighted tensions between safety advocacy and corporate confidentiality.
Legislative progress on the FRONTIER Act and similar state-level initiatives in California will be key indicators of whether the government will move to standardize the role of Independent Verification Organizations (IVOs).
The industry will monitor whether the 'Minimum Conditions for Evaluators' proposed by the AI Evaluator Forum gains traction as a de facto standard for how these organizations operate within private labs.
Future reports from these evaluators will be scrutinized for their depth and independence. If auditors are unable to report unfavorable findings without fear of losing their contracts, the credibility of the entire third-party evaluation ecosystem may be undermined.