Top storyAnthropic Says Fable 5 Biology Fallbacks Fell 85%
Anthropic says a retrained safety classifier reduced biology-related fallbacks by about 85%, allowing more health and education queries while keeping dual-use research restricted.
Updated daily124 verified stories
AI safety, security, alignment, evaluations, incidents, vulnerabilities, and safeguards tracked without hype.
Every story links to the strongest available evidence: original sources when available, otherwise clearly attributed reporting.
What happened, why it matters, and what to watch — no jargon tax.
When the signal is thin, we publish nothing rather than padding the feed.
A growing stream of verified perspectives for people who need to understand AI without chasing hype.
Top storyAnthropic says a retrained safety classifier reduced biology-related fallbacks by about 85%, allowing more health and education queries while keeping dual-use research restricted.

A six-model preprint found sharply different ways of resisting, redirecting, or following steering prompts, but its model-judge labels still need validation by human raters.

A UK AI Security Institute-affiliated preprint reproduced several safety-benchmark results with 97–99% fewer prompts, while warning that shorter tests do not prove real-world safety.
One useful briefing each week
Get the week’s verified AI news, original data, useful tools, learning picks, and fresh AI jobs.
Hiring an AI professional or launching a useful AI product? Put it in front of people who came here to learn and act.
Post an AI job Submit an AI tool