Back to News
PolicyAI Understanding briefing

Scale AI calls for increased federal funding for frontier model testing

Scale AI argues that current U.S. government funding for AI evaluation is insufficient to address national security risks, citing new research on agent vulnerabilities.

4 min readRead the primary source
Source-provided image accompanying Scale AI calls for increased federal funding for frontier model testing
Primary-source documentSource recorded
Publisher
scale.com
Source link
scale.comhttps://scale.com/blog/steer-the-ai-frontier-washington-must-build-its-testing-power
Source type
Primary document — an official announcement, paper, filing, or first-party page we read directly.
ContextUnderstand this in 60 seconds

Start here

Test yourselfAI Ethics Quiz

What happened

Scale AI has issued a public call for the U.S. government to significantly increase funding and infrastructure for independent frontier AI model testing. The company argues that current federal investment levels—specifically the $15 million to $27 million proposed for the government's frontier AI testing center (CAISI) in fiscal year 2027—are inadequate for the scale of current AI development. Scale AI asserts that the government must move beyond relying on internal lab testing to establish its own evidence-based benchmarks for cyber, biological, and autonomy-related risks.

Scale AI’s position highlights a growing gap between the rapid advancement of frontier AI models and the federal government's capacity to evaluate them. The company argues that the current proposed funding for the government's frontier AI testing center, CAISI, is insufficient to meet the demands of modern AI oversight.

The company emphasizes that the government must establish its own independent testing infrastructure to assess risks in areas such as cyber defense, biological threats, and autonomous systems. This would shift the burden of proof away from private labs, which currently perform the majority of safety testing internally.

Source details: scale.com ↗

Why it matters

The debate over AI regulation is currently hindered by a lack of standardized, independent measurement. By relying on internal company testing, policymakers risk either over-regulating based on speculation or under-regulating critical security threats. Scale AI’s findings suggest that current models possess specific, overlooked vulnerabilities—such as susceptibility to malicious instructions during task-clarification pauses and a tendency to refuse assistance in defensive cyber scenarios—that could jeopardize national infrastructure like power grids and banking systems if not properly evaluated by independent, government-backed entities.

The core issue is the 'diagnostic gap' in AI policy. Without independent, standardized testing, policymakers cannot distinguish between hypothetical risks and demonstrated threats. This uncertainty complicates the development of effective regulations that protect public safety without stifling innovation.

Scale AI’s internal research provides concrete examples of why this matters: they found that AI agents are vulnerable to 'instruction injection' when they pause to ask for clarification during tasks like banking or email management. Furthermore, they observed that models often refuse to assist in defensive cyber operations, which could leave critical infrastructure vulnerable during real-world attacks.

These findings underscore the necessity for testing environments that simulate real-world threats rather than relying on standard, static benchmarks that may miss nuanced security failures.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Interactive Concept Check+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

What to watch next

Observers should monitor upcoming federal budget negotiations regarding the CAISI testing center and potential new partnerships between the government and private evaluation firms. Scale AI has indicated it will share more information regarding its ongoing work with international government evaluation bodies in the coming weeks. Additionally, the effectiveness of future policy will depend on whether the government can successfully transition from theoretical risk assessment to the implementation of rigorous, standardized diagnostic benchmarks for frontier models.

The primary focus is on the legislative process surrounding the fiscal year 2027 budget for AI evaluation. The discrepancy between the President's $27 million request and the House's $15 million proposal remains a key point of contention.

Scale AI has promised to release more information regarding its collaborations with government bodies in the U.S., U.K., Korea, and Singapore. These partnerships may serve as a blueprint for how the private sector can support government-led evaluation efforts.

The industry will be watching to see if the government adopts a more aggressive stance on independent testing, which would fundamentally change the compliance landscape for companies developing frontier models.

Related guides & quizzes

AI EthicsAI Models ExplainedAI AgentsFuture of AITest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI regulation tracker
Found this useful?