Kini o ṣẹlẹ
Anthropic imudojuiwọn Claude Awọn aabo isedale itan-akọọlẹ 5 ni Oṣu Kẹjọ Ọjọ 7, titọka ikasi kan ti o ti yi pada fere gbogbo ibeere isedale si awoṣe ti o ni agbara biology.
Nigbati olutọpa naa ba beere ibeere kan, Anthropic ṣe ọna rẹ lati Fable 5 si Opus 5. Ile-iṣẹ sọ pe Opus 5 wa ni agbara fun lilo gbogbogbo ṣugbọn pese iranlọwọ iṣẹ ti o dinku lori isedale ti ilọsiwaju, dinku iye eto naa si ẹnikan ti o lepa iṣẹ ipalara.
Anthropic sọ pe o ṣe atunto ofin olupilẹṣẹ, kojọ awọn esi lati ọdọ awọn amoye inu ati ita, ṣẹda data ikẹkọ tuntun, ṣe atunṣe , ati ṣayẹwo pe o tun fa ni gbogbogbo lori ipalara ati awọn ibeere iwadii ilo-meji. Ninu idanwo ile-iṣẹ naa, imudojuiwọn naa dinku awọn ifẹhinti ti o ni ibatan isedale nipasẹ iwọn 85% kọja awọn ọja rẹ.
The change follows a deliberately conservative launch posture. Anthropic says Fable 5 initially routed almost every biology request to Opus 5 because the company preferred a broad safety boundary while it learned where benign health and education questions were being caught. The August update narrows that boundary through a separate instead of changing the model's underlying biology capability. That makes the release a policy-and-routing change, not evidence that Fable 5 has become a clinically validated biology assistant.
Iyipada naa jẹ itumọ lati jẹ ki Fable 5 dahun ilera lojoojumọ diẹ sii, ile-iwosan, ati awọn ibeere eto-ẹkọ. Anthropic sọ pe wiwọle lasan tun ṣubu pada fun awọn agbegbe lilo-meji pẹlu virology, toxicology, ati apẹrẹ molikula, nitorinaa awoṣe ko tii wa nipasẹ ọna yẹn fun iwadii isedale alamọdaju tabi idagbasoke oogun.
Awọn alaye orisun: Anthropic's Fable 5 biology safeguards announcement ↗
Kini idi ti o ṣe pataki
Imudojuiwọn naa jẹ idanwo ti o wulo ti boya awọn aabo awoṣe-aala le di kongẹ diẹ sii laisi yiyan larọrun laarin iraye si gbooro ati kiko gbooro.
Àlẹmọ isokuso le dinku eewu ni kiakia, ṣugbọn o tun le dina awọn ọmọ ile-iwe, awọn alaisan, awọn olukọni, ati awọn alamọdaju ilera ti awọn ibeere wọn lo ede imọ-ẹrọ kanna bi iwadii ifura. Anthropic yan aaye ibẹrẹ Konsafetifu yẹn nigbati o ṣe ifilọlẹ Fable 5, lẹhinna lo eto imulo alaye diẹ sii ati awọn apẹẹrẹ ikẹkọ tuntun lati gbe aala fun awọn ibeere alaiwu.
Iyipada iriri olumulo yatọ nipasẹ ọja nitori isedale jẹ idi kan nikan ti ipadasẹhin. Anthropic ṣe iṣiro pe lapapọ awọn ipadasẹhin ti gbogbo iru yoo ṣubu nipasẹ aijọju 67% lori Claude.ai, 55% ni Iṣẹ-iṣẹ, 17% ni Claude Koodu, ati 7% lori Claude Platform. Iyẹn jẹ awọn wiwọn ile-iṣẹ, kii ṣe awọn abajade iṣayẹwo ominira.
Ilana naa tun ṣe pataki: ibeere ti a fi ami si ni a tun pada ju ki o dahun nipasẹ Fable 5. Ti o tọju iraye si awoṣe gbogbogbo lakoko ti o dawọ agbara Anthropic ṣe akiyesi pupọ julọ, ṣugbọn ko fi idi rẹ mulẹ pe gbogbo idahun ilera ti a gba laaye jẹ deede tabi yẹ fun ipinnu ile-iwosan.
For organizations, the practical question is how the boundary behaves across contexts. A student asking for a plain-language explanation, a clinician checking terminology, and a researcher requesting an experimental protocol may use overlapping words while presenting very different risk. Anthropic's routing approach can preserve a safer general answer for the first two cases, but only if the recognizes intent, conversation history, and requested operational detail without turning a legitimate professional workflow into an opaque denial.
Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ
Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.
crm_get_transaction(id='4092').What is 'specification gaming' in AI systems?
Kini lati wo tókàn
Ṣọra fun ẹri pe iwọn-pada sẹhin ni ibamu nipasẹ wiwa to lagbara ti awọn ibeere ti o lewu nitootọ, pẹlu awọn ofin mimọ fun iraye si iwadii igbẹkẹle.
Anthropic ko ṣe atẹjade eto igbelewọn, oṣuwọn odi eke, tabi ẹda ominira pẹlu ikede yii. Ile-iṣẹ sọ pe awọn idaniloju eke yoo wa ati pe awọn kilasika gbọdọ tun koju awọn igbiyanju isakurolewon, nitorinaa eeya 85% ṣe iwọn awọn apadabọ diẹ kuku ju iṣowo ailewu ni kikun.
Awọn iwifun ti o wulo ti o tẹle yoo ṣe afihan iṣẹ ṣiṣe kọja awọn atupalẹ, awọn ede, awọn ibaraẹnisọrọ ọpọlọpọ-iyipada, ati ṣiṣiṣẹsiṣẹ-ṣiṣẹ irinṣẹ. Awọn oniwadi tun nilo lati mọ iye igba ti ibeere ipalara kan kọja aala tuntun ati bawo ni iyara ti ṣe imudojuiwọn kilasika nigbati awọn ipadabọ tuntun han.
Anthropic sọ pe o n ṣe idagbasoke awọn ọna iraye si igbẹkẹle fun awọn agbara isedale iwaju. Igbẹkẹle wọn yoo dale lori ẹniti o yẹ, kini ibojuwo ati awọn aabo ikọkọ ti o waye, bawo ni a ṣe n ṣe atunyẹwo awọn iṣẹlẹ, ati boya awọn oniwadi to tọ le koju ihamọ ti ko tọ.
The next disclosure should also explain how the is evaluated after deployment. A lower fallback rate can be achieved by reducing false positives, by shifting difficult cases to another model, or by missing more harmful requests; those outcomes have very different safety meanings. Useful reporting would include false-positive and false-negative estimates by request type, performance under multi-turn escalation and paraphrase, language coverage, handling of tool calls, and the process for updating the boundary after a jailbreak or an incident.
The user-facing promise should be tested with the same care as the safety boundary. Anthropic says the update should help with everyday health, clinical, and educational questions, but a lower fallback rate does not establish medical accuracy, appropriate triage, or suitability for professional decisions. Independent reviewers should sample allowed answers for unsupported certainty, missing safety advice, and harmful procedural detail, while also checking whether the fallback model communicates its limits clearly. The strongest evidence would compare matched requests before and after the change and publish enough anonymized examples for outside researchers to understand both the gains and the new failure modes.
That evidence should be reported separately for consumer chat, coding, agentic workflows, and the API because the same boundary may carry different tools, context windows, and user expectations in each product. A single blended fallback percentage can hide a meaningful regression in one surface behind improvement in another. Publishing the denominator, confidence intervals, and product-level counts would make the result useful to educators, clinicians, developers, and safety researchers rather than only to readers comparing one headline number.