Pada si Iroyin
ÀàbòAI Understanding finifini

Awọn oludasilẹ ṣe ijabọ OpenAI ati awọn aabo Anthropic fa fifalẹ iṣẹ ṣiṣe iranlọwọ AI ni igbagbogbo

Ni Ọjọ Dev ti OpenAI, awọn olupilẹṣẹ sọ pe awọn oluso aabo awoṣe n ṣe afihan afẹfẹ afẹfẹ lasan, awọn roboti ati awọn iṣẹ-ṣiṣe cybersecurity bi eewu, fi ipa mu wọn lati yi awọn awoṣe pada tabi fi awọn itọsi silẹ, ati ṣafikun awọn idiyele akoko ti o farapamọ si ṣiṣan iṣẹ wọn.

4 min readRead the original reporting
Source-provided image accompanying Developers report OpenAI and Anthropic safeguards slowing routine AI‑assisted work
Ijabọ iroyinOrisun ti o gbasilẹ
Olutẹwe
venturebeat.com
Orisun ọna asopọ
venturebeat.comhttps://venturebeat.com/technology/developers-say-openai-and-anthropic-safeguards-are-flagging-routine-work-and-costing-them-time
Orisun iru
Ijabọ nipasẹ ijade iroyin kan - kii ṣe iwe-ipamọ ẹgbẹ akọkọ.

Ohun ti a ko le jẹrisi ni ominira: Ibeere yii jẹ ikasi si iṣan ti a npè ni. A ko jẹrisi rẹ lodi si iwe-ipamọ ẹgbẹ akọkọ. (venturebeat.com)

AtokọLoye eyi ni iṣẹju 60

Bẹrẹ nibi

Awọn ofin bọtini

Isọdiwọn
Bii awọn iṣiro igbẹkẹle awoṣe ṣe baamu awọn iṣeeṣe deede gangan.
Awọn ọna opopona
Awọn ofin, sọwedowo, ati awọn idari ti o fi opin si ailewu tabi ihuwasi awoṣe aifẹ.
AI Aabo
Aaye kan lojutu lori idinku ihuwasi ipalara, awọn ikuna, ati awọn ewu ilokulo ninu awọn eto AI.
Ṣe idanwo fun ara rẹAI Ethics adanwo

Kini o ṣẹlẹ

Developers told VentureBeat that OpenAI’s newly announced GPT‑6.1 Sol and Anthropic’s Claude models are increasingly flagging routine prompts as unsafe, even when the requests are harmless. Interviewees described frequent “refusals” when asking models to access device parameters, generate user interfaces for robot arms, or discuss cybersecurity topics. Some users abandoned the chat, moved tasks to older models, or switched to open‑source alternatives such as Moonshot AI’s Kimi or Alibaba’s Qwen. OpenAI acknowledged that its safeguards can “slow, pause, or stop legitimate work,” and said it is refining the . Anthropic, after receiving complaints about its Fable 5 series, released Fable 5.1 with a claimed 60 % reduction in interventions for code‑related queries, though certain penetration‑testing prompts still route away. The article also cites JetBrains’ 2026 Developer Ecosystem Survey, which found 90 % of 15,000 surveyed developers use AI coding agents weekly, underscoring the breadth of impact.

During OpenAI’s Dev Day keynote, the company announced GPT‑6.1 Sol, a model positioned as a cheaper alternative to GPT‑6 Astra for coding and computer‑use tasks. The same announcement noted that the model retains the same safety stack as Astra, which is classified as "Critical" under OpenAI’s Preparedness Framework for cybersecurity capabilities.

Developers interviewed by VentureBeat reported that both OpenAI’s and Anthropic’s models are increasingly flagging routine prompts—such as generating UI code for robotic arms, SSH connections, or discussing satellite simulations—as risky. The refusals often appear after a series of benign queries, making it hard to reset the conversation.

OpenAI confirmed to the reporter that its safeguards can interrupt legitimate work, even in defensive cybersecurity contexts, and said it is working to reduce unnecessary interruptions. Anthropic, after acknowledging false‑positive issues with its Fable 5 series, released Fable 5.1, claiming a 60 % drop in interventions for code‑related sessions, though certain penetration‑testing queries still trigger redirects.

Developers are responding by switching to older models, open‑source alternatives, or locally‑run LLMs that lack the same . Some have canceled subscriptions to more restrictive services, while others rely on OpenAI’s Daybreak Access program for qualified enterprise customers to obtain more permissive access.

Awọn alaye orisun: venturebeat.com ↗

Kini idi ti o ṣe pataki

The growing friction between safety mechanisms and productive developer workflows highlights a tension at the core of AI deployment: protecting users and systems without crippling legitimate use cases. If safeguards routinely block benign requests, developers may incur hidden time costs, delay product timelines, or abandon higher‑performing frontier models in favor of less capable or open‑source alternatives. This could slow adoption of advanced AI tools in high‑stakes domains such as aerospace, robotics, and cybersecurity, where rapid iteration is critical. Moreover, the reported false‑positive rates raise questions about the of and the transparency of safety policies, especially as OpenAI’s GPT‑6 line claims “critical” cybersecurity capabilities. The situation also illustrates the market pressure on AI vendors to balance safety with usability, influencing pricing, access programs like OpenAI’s Daybreak Access, and the competitive dynamics between closed and open models.

The friction caused by over‑zealous safeguards can translate into hidden productivity costs for developers, especially in sectors where rapid prototyping and iteration are essential. This may slow the broader adoption of cutting‑edge AI assistants in critical industries.

Safety mechanisms that generate false positives undermine trust in AI systems, potentially prompting developers to favor less capable or open‑source models that lack robust , thereby affecting market dynamics and revenue for leading AI firms.

OpenAI’s claim that GPT‑6 Astra can discover and exploit unknown vulnerabilities without step‑by‑step human direction raises stakes for how such powerful models are governed. If safeguards impede legitimate defensive work, security teams may be forced to operate without the most advanced tools, affecting overall cyber resilience.

Interactive Mechanism

Ibaraẹnisọrọ Mechanism: Bii O Ṣe Nṣiṣẹ Lootọ

Ṣawari imọ-ẹrọ abẹlẹ lẹhin idagbasoke yii ni ibaraenisọrọ.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Ibanisọrọ Erongba Ṣayẹwo+10 Points
AI Ethics Quiz

Impossibility results in algorithmic fairness (e.g. Kleinberg et al., Chouldechova) show what?

Kini lati wo tókàn

Future updates from OpenAI and Anthropic on the effectiveness of revised , especially any quantitative metrics on reduced false positives. Adoption rates of Daybreak Access or similar trusted‑access programs among enterprise cybersecurity teams. Potential shifts toward open‑source or locally‑run models if closed‑source safeguards remain disruptive. Industry responses, such as new developer‑focused tooling that can pre‑filter or re‑phrase prompts to avoid refusals, and any regulatory scrutiny of practices that affect productivity.

Metrics from OpenAI and Anthropic on the frequency of false‑positive refusals after the rollout of GPT‑6.1 Sol and Fable 5.1, respectively.

Adoption trends for Daybreak Access and similar trusted‑access programs, including any changes to eligibility criteria or pricing.

Emergence of third‑party tooling designed to pre‑process prompts to avoid guardrail triggers, and whether such tools gain traction among developer communities.

Regulatory developments concerning standards that could mandate transparency or limit the aggressiveness of model safeguards.

Awọn itọsọna ti o jọmọ & awọn ibeere

Ìlànà Ìwà AIAwọn aṣoju AIAwọn awoṣe AI ti ṣalayeṢe idanwo ohun ti o mọ — gbiyanju idanwo AI ọfẹ kanWa ọrọ AI kan ninu iwe-itumọ waTẹle olutọpa ilana AI
Ṣe eyi wulo?