Quay lại Tin tức
Bảo mậtAI Understanding tóm tắt

Capsule Security báo cáo bộ ngắt mạch AI dựa trên Nemotron

SiliconANGLE báo cáo rằng Capsule Security đã phát hành một hệ thống dựa trên Nemotron để đánh giá các hành động của tác nhân AI ngay trước khi thực thi và có thể cho phép, gắn cờ hoặc chặn chúng.

4 min readRead the linked source
Source-provided image accompanying Capsule Security reports Nemotron-based AI circuit breaker
Nguồn tham khảoNguồn đã ghi
Nhà xuất bản
siliconangle.com
Liên kết nguồn
siliconangle.comhttps://siliconangle.com/2026/09/02/capsule-security-fine-tunes-nvidia-nemotron-models-to-stop-rogue-ai-agents/
Loại nguồn
Nguồn được liên kết - trạng thái nguồn chính chưa được thiết lập.
Cũng được trích dẫn

Câu chuyện được sửa đổi lần cuối

Bối cảnhHiểu điều này trong 60 giây

Bắt đầu ở đây

Thuật ngữ chính

Bộ nhớ (Bộ nhớ tác nhân)
Bối cảnh được lưu trữ mà tác nhân AI sử dụng qua các bước hoặc phiên để cải thiện tính liên tục.
tiêm nhắc nhở
Một kiểu tấn công trong đó các lệnh độc hại được chèn vào đầu vào của mô hình hoặc nội dung được truy xuất.
Điểm chuẩn
Một bài kiểm tra hoặc tập dữ liệu được tiêu chuẩn hóa dùng để đo lường và so sánh hiệu suất của mô hình.
Tự kiểm traCâu đố về đại lý AI

Điều gì đã thay đổi kể từ khi xuất bản

  1. Xuất bản lần đầu
  2. SiliconANGLE adds that Capsule Security has released a separate Nemotron-based pre-execution monitor for rogue AI-agent actions. Capsule reports 98% accuracy on StepShield, 96.9% on an internal benchmark, response times as low as 71 milliseconds and single-NVIDIA-L40S deployment for its larger model. The capability is said to be available now, but pricing, access conditions, integrations and independent validation are unknown.

Chuyện gì đã xảy ra

SiliconANGLE reports that Capsule Security released an “AI circuit breaker” built from two fine-tuned NVIDIA Nemotron models. The system evaluates an agent’s intended action immediately before execution and lets customers allow, flag or block it. Capsule said the detector reached 98% accuracy on the StepShield and returned decisions in as little as 71 milliseconds. The capability is reportedly available now, although access conditions and pricing were not disclosed.

SiliconANGLE reports that Capsule Security released a detection system based on two NVIDIA Nemotron models that it fine-tuned itself. Capsule describes the product as an “AI circuit breaker”: it judges an agent’s intended action immediately before the action executes, allowing a customer to permit, flag or block it. The source says the control layer is intended for agents with credentials to sensitive data, source code or production infrastructure.

According to SiliconANGLE, Capsule said its system achieved 98% accuracy on StepShield, a using 9,429 code-agent trajectories drawn from real incidents. Capsule’s most accurate detector reportedly scored 96.9% on an internal benchmark, compared with 86% for the strongest unnamed third-party model tested. The article says Capsule did not identify that model or provide a score breakdown.

SiliconANGLE reports that decisions returned in as little as 71 milliseconds. Capsule said its larger model’s memory requirements were reduced by nearly half without a performance loss, allowing it to run on one NVIDIA L40S GPU. The training data reportedly included real agent traces and human-reviewed adversarial examples marking the boundary between authorized and unauthorized behavior.

The source says the capability is available now. It does not specify pricing, procurement, supported agent platforms, geographic limits or whether the release is generally accessible. None of the performance claims, customer references or the reported availability was independently confirmed for this evaluation.

Chi tiết nguồn: siliconangle.com ↗

Tại sao nó quan trọng

The reported system targets a practical weakness in agent security: permissions can limit what an agent may access, but they do not necessarily determine whether a specific action is appropriate for the task. A pre-execution decision layer could give security teams another control point before an error becomes an incident. The performance figures, customer claims and comparisons remain SiliconANGLE’s account of Capsule’s statements and were not independently confirmed here.

The reported approach focuses on whether an action fits the task an agent was assigned, rather than only checking whether the agent technically has permission to perform it. That distinction matters as organizations give agents access to credentials, code repositories and production systems. A control applied before execution could limit the time between an unsafe decision and a possible consequence, though the source provides no independent evidence that it prevents real-world incidents.

The reported latency and single-GPU deployment claim could make step-level monitoring more practical for some organizations than a large generative model used as a separate reviewer. However, the article does not provide false-positive or false-negative rates, examples of blocked actions, testing methodology beyond the cited , or independent replication. The comparison with unnamed third-party and frontier models therefore cannot establish superiority.

SiliconANGLE says customers include financial institutions and technology companies and quotes H&R Block’s chief security information officer supporting controls of this kind. Those references indicate claimed enterprise interest, not proof of broad deployment or effectiveness. The source does not state whether the product is self-serve, enterprise-only or sold as a separate paid service.

Interactive Mechanism

Cơ chế tương tác: Nó thực sự hoạt động như thế nào

Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Kiểm tra khái niệm tương tác+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Xem gì tiếp theo

Watch for independent testing of Capsule’s StepShield results, details on false positives and false negatives, and evidence from production deployments. It is also unclear how the system integrates with different agent frameworks, what actions it can inspect, whether human approval is required for blocked or flagged actions, and whether the reported 71-millisecond latency holds under enterprise workloads. SiliconANGLE did not report pricing or a public self-serve sign-up path.

Independent researchers and customers should test whether the reported accuracy transfers from StepShield and Capsule’s internal to different agent frameworks, tools, tasks and attack methods. Step-level accuracy alone may not show how often a monitor misses a harmful action or interrupts legitimate work.

Further reporting should clarify how Capsule handles ambiguous actions, cascading tool calls, , compromised credentials and agents whose behavior changes after monitoring. The source does not explain what evidence the detector uses to decide whether an action is authorized or how customers configure policies.

The capability is described as available now, but SiliconANGLE provides no price, sign-up process, service-level terms or list of supported integrations. Those details will determine who can actually use it and whether the product is practical for smaller organizations as well as the enterprise customers mentioned in the report.

Hướng dẫn và câu hỏi liên quan

Đại lý AIGiải thích về mô hình AIĐạo đức AIKiểm tra những gì bạn biết — thử một bài kiểm tra AI miễn phíTra cứu một thuật ngữ AI trong bảng thuật ngữ của chúng tôiThực hiện theo trình theo dõi quy định AI

Cập nhật và sửa chữa

Câu chuyện kinh điển này được cập nhật tại chỗ khi sự kiện đang phát triển có thay đổi cơ bản. URL và ngày xuất bản ban đầu của nó không bao giờ thay đổi.

  • SiliconANGLE adds that Capsule Security has released a separate Nemotron-based pre-execution monitor for rogue AI-agent actions. Capsule reports 98% accuracy on StepShield, 96.9% on an internal benchmark, response times as low as 71 milliseconds and single-NVIDIA-L40S deployment for its larger model. The capability is said to be available now, but pricing, access conditions, integrations and independent validation are unknown.
Xem nhật ký chỉnh sửa công khai
Tìm thấy điều này hữu ích?