Back to News
PolicyAI Understanding briefing

AI assessors say current science lags behind safety measures people want

Anthropic and OpenAI announced they are working with third‑party evaluators to certify their models, but several assessors warned that the scientific methods used to gauge AI safety have not yet caught up with the expectations of regulators and the public.

4 min readRead the original reporting
Source-provided image accompanying AI assessors say current science lags behind safety measures people want
Attributed reportingSource recorded
Publisher
npr.org
Source link
npr.orghttps://www.npr.org/2026/09/28/nx-s1-5974206/ai-assessors-says-current-science-hasnt-caught-up-to-the-safety-measures-people-want
Source type
Reporting by a news outlet — not a first-party document.

What we could not confirm independently: This claim is attributed to the named outlet. We did not verify it against a first-party document. (npr.org)

ContextUnderstand this in 60 seconds

Start here

Key terms

Generative AI
AI systems that produce new content such as text, images, audio, video, or code.
Robustness
A model's ability to maintain performance under noise, shifts, or adversarial inputs.
AI Safety
A field focused on reducing harmful behavior, failures, and misuse risks in AI systems.
Test yourselfAI Ethics Quiz

What happened

Anthropic and OpenAI told NPR they are collaborating with independent safety assessors to evaluate the risks of their latest AI systems. The assessors, whose identities were not disclosed, cautioned that existing science—such as testing, interpretability research, and alignment metrics—lags behind the safety standards that policymakers, industry groups, and the public are demanding. The companies said the partnership aims to improve transparency and address growing concerns about AI‑driven harms, but the evaluators highlighted gaps in current methodologies, noting that many safety measures remain experimental and lack standardized benchmarks.

In a September 28, 2026 interview with NPR, representatives from Anthropic and OpenAI said they are actively engaging third‑party evaluators to assess the safety of their latest offerings. The companies emphasized a commitment to transparency and to addressing public concerns about potential harms such as misinformation, bias, and unintended behavior.

The evaluators, speaking on condition of anonymity, warned that the current scientific toolkit for is still catching up. They pointed to a lack of universally accepted metrics for alignment, limited reproducibility of tests, and insufficient data on long‑term model behavior. These gaps mean that safety certifications may rely on provisional methods that have not yet been validated across diverse real‑world scenarios.

Both firms indicated that the partnership with assessors is part of a broader effort to develop more rigorous testing regimes, but they did not provide specific timelines or details about the methodologies being employed. The conversation highlighted a tension between the speed of AI product releases and the slower pace of safety research.

Source details: npr.org ↗

Why it matters

The statements underscore a widening gap between the rapid deployment of powerful AI models and the development of rigorous, peer‑reviewed safety science. As AI systems become more capable, regulators worldwide are considering stricter oversight, and industry leaders are under pressure to demonstrate that their products meet high safety thresholds. If scientific tools cannot keep pace, there is a risk that safety certifications become symbolic rather than substantive, potentially allowing unsafe models to reach users. The gap also fuels legislative proposals, such as Senator Warner’s mandatory pre‑release testing bill, and could influence state‑level actions like Oregon’s AI safety regulations. Understanding where the science falls short helps policymakers target funding, prioritize research, and set realistic timelines for regulatory frameworks.

The gap between safety expectations and scientific capability raises the stakes for regulators who must decide whether to impose restrictions on AI deployment. Without robust, peer‑reviewed evidence, policy decisions risk being based on incomplete or anecdotal data.

Industry credibility is at risk. If major AI firms cannot demonstrate that their safety assessments are grounded in solid science, public trust may erode, potentially prompting consumer backlash or stricter market controls.

The statements may accelerate funding and research initiatives aimed at closing the safety science gap, influencing academic and private‑sector priorities. This could lead to new collaborations, standards bodies, and possibly the creation of dedicated safety research institutes.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

System Requirements:
Best ArchitecturePure RAGRecommended pattern
Hallucination RiskVery LowGrounding efficacy
Update Cost$0 (Vector sync)Ongoing maintenance
Core takeaway: Fine-tuning teaches models how to speak (form, style, syntax); RAG teaches models what to say (verifiable facts). Never use fine-tuning alone for factual memory.
Interactive Concept Check+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

What to watch next

Future disclosures from Anthropic, OpenAI, and other AI firms about concrete safety testing protocols; any new standards or guidelines issued by bodies such as NIST or the International Organization for Standardization; legislative activity at the federal and state level that references the need for stronger scientific foundations; and the emergence of third‑party organizations that aim to create standardized safety benchmarks for AI models.

Announcements of concrete safety testing frameworks from Anthropic, OpenAI, or other leading AI developers, especially those that reference standardized benchmarks or third‑party certification processes.

Legislative proposals that cite the need for stronger scientific foundations, such as bills mandating pre‑release safety evaluations or requiring public disclosure of safety testing results.

The formation of new independent safety assessment organizations or the expansion of existing ones, which could provide the missing scientific rigor and serve as a bridge between industry and regulators.

International standard‑setting efforts, for example by NIST or ISO, that aim to codify metrics and testing protocols.

Related guides & quizzes

AI EthicsFuture of AIWhat is AI?Test what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI regulation tracker
Found this useful?