发生了什么
Anthropic 于 8 月 7 日更新了 Claude 《神鬼寓言 5》的生物学保护措施,缩小了分类器的范围,该分类器已将几乎所有生物学查询重新路由到生物学能力较低的模型。
当分类器标记一个请求时,Anthropic 将其从 Fable 5 路由到 Opus 5。该公司表示,Opus 5 仍然具有一般用途的能力,但对高级生物学提供的操作帮助较少,从而降低了系统对于从事有害工作的人的价值。
Anthropic 表示,它重写了分类器的构成,收集了内部和外部专家的反馈,创建了新的训练数据,重新训练了分类器,并检查它是否仍然普遍触发有害和双重用途的研究请求。在该公司的测试中,该更新将其产品中与生物学相关的故障减少了约 85%。
The change follows a deliberately conservative launch posture. Anthropic says Fable 5 initially routed almost every biology request to Opus 5 because the company preferred a broad safety boundary while it learned where benign health and education questions were being caught. The August update narrows that boundary through a separate instead of changing the model's underlying biology capability. That makes the release a policy-and-routing change, not evidence that Fable 5 has become a clinically validated biology assistant.
这一变化旨在让《神鬼寓言 5》回答更多日常健康、临床和教育问题。 Anthropic 表示,对于包括病毒学、毒理学和分子设计在内的双重用途领域,普通访问仍然落后,因此该模型尚无法通过该路径用于专业生物学研究或药物开发。
来源详情: Anthropic's Fable 5 biology safeguards announcement ↗
为什么这很重要
此次更新是对前沿模型保障措施是否可以变得更加精确而无需简单地在广泛准入和广泛拒绝之间进行选择的实际测试。
粗略过滤可以快速降低风险,但也会阻止学生、患者、教育工作者和医疗保健专业人员提出与敏感研究相同的技术语言。 Anthropic 在发布 Fable 5 时选择了保守的起点,然后使用更详细的策略和新的训练示例来移动良性请求的边界。
用户体验的变化因产品而异,因为生物学只是后退的原因之一。 Anthropic 估计 Claude.ai 上的各种总回退将下降约 67%,Cowork 中下降 55%,Claude 代码中下降 17%,Claude 平台上下降 7%。这些是公司的衡量结果,而不是独立的审计结果。
该机制也很重要:标记的请求会被重新路由,而不是由寓言 5 应答。这保留了对通用模型的访问,同时保留了 Anthropic 认为最令人担忧的功能,但它并不能确定每个允许的健康答案都是准确的或适合临床决策的。
For organizations, the practical question is how the boundary behaves across contexts. A student asking for a plain-language explanation, a clinician checking terminology, and a researcher requesting an experimental protocol may use overlapping words while presenting very different risk. Anthropic's routing approach can preserve a safer general answer for the first two cases, but only if the recognizes intent, conversation history, and requested operational detail without turning a legitimate professional workflow into an opaque denial.
互动机制:它实际上是如何运作的
以交互方式探索这一发展背后的基础技术。
crm_get_transaction(id='4092').What is 'specification gaming' in AI systems?
接下来看什么
留意证据表明,较低的回退率与对真正危险请求的强大检测相匹配,以及可信研究访问的明确规则。
Anthropic 没有随本公告发布评估集、假阴性率或独立复制。该公司表示,误报仍然存在,分类器还必须能够承受越狱尝试,因此 85% 的数字衡量的是更少的后备,而不是完整的安全权衡。
下一个有用的披露将显示跨释义、语言、多轮对话和支持工具的工作流程的性能。研究人员还需要知道有害请求跨越新边界的频率以及当新的绕过出现时分类器的更新速度。
Anthropic 表示正在为前沿生物学能力开发可信访问途径。它们的可信度将取决于谁有资格、适用哪些监控和隐私保护、如何审查事件以及合法的研究人员是否可以挑战不正确的限制。
The next disclosure should also explain how the is evaluated after deployment. A lower fallback rate can be achieved by reducing false positives, by shifting difficult cases to another model, or by missing more harmful requests; those outcomes have very different safety meanings. Useful reporting would include false-positive and false-negative estimates by request type, performance under multi-turn escalation and paraphrase, language coverage, handling of tool calls, and the process for updating the boundary after a jailbreak or an incident.
The user-facing promise should be tested with the same care as the safety boundary. Anthropic says the update should help with everyday health, clinical, and educational questions, but a lower fallback rate does not establish medical accuracy, appropriate triage, or suitability for professional decisions. Independent reviewers should sample allowed answers for unsupported certainty, missing safety advice, and harmful procedural detail, while also checking whether the fallback model communicates its limits clearly. The strongest evidence would compare matched requests before and after the change and publish enough anonymized examples for outside researchers to understand both the gains and the new failure modes.
That evidence should be reported separately for consumer chat, coding, agentic workflows, and the API because the same boundary may carry different tools, context windows, and user expectations in each product. A single blended fallback percentage can hide a meaningful regression in one surface behind improvement in another. Publishing the denominator, confidence intervals, and product-level counts would make the result useful to educators, clinicians, developers, and safety researchers rather than only to readers comparing one headline number.