返回新闻
安全AI Understanding 简报

AWS 基准测试发现 AI 漏洞检测器标记安全代码

Help Net Security 报告称,AWS 在基准测试中测试了 12 个人工智能模型,该基准测试旨在区分可利用的漏洞和看起来危险的安全代码。没有一个达到 AWS 规定的假阳性和假阴性率生产阈值。

4 min readRead the linked source
Source-provided image accompanying AWS benchmark finds AI vulnerability detectors flag safe code
来源参考来源记录
出版商
helpnetsecurity.com
来源链接
helpnetsecurity.comhttps://www.helpnetsecurity.com/2026/09/14/aws-deception-benchmark-security-vulnerabilities/
来源类型
链接来源——主要来源状态尚未确定。
背景60 秒内了解这一点

从这里开始

关键术语

基准测试
用于测量和比较模型性能的标准化测试或数据集。
精度
实际正确的预测阳性的比例。
测试一下自己AI 模型解释测验

发生了什么

Help Net Security reports that AWS has publicly released its Deception , which tests whether AI models can distinguish genuine software vulnerabilities from safe code containing misleading vulnerability patterns. The benchmark includes 14,822 samples spanning 16 programming languages and more than 70 CWE categories; 9,695 are scored and 5,127 are mixed in as unscored samples.

Help Net Security reports that AWS evaluated 12 general-purpose models from five providers using a designed around false positives. The benchmark’s safe examples contain real vulnerability patterns alongside protections that prevent exploitation. Some challenges vary deployment conditions, such as Kubernetes network policies, so a model must consider both the code and its environment.

According to the report, AWS generated and refined examples against frontier models, excluding samples that were too easy to classify. AWS says this process consumed tens of billions of tokens. The company publicly releases the samples but withholds their labels; users submit predictions to AWS for verified scoring. The source does not state whether access to scoring carries a fee or other eligibility requirements.

Help Net Security reports that independent reviewers check labels without seeing one another’s decisions or the original reasoning. Disputed samples receive further review, and unresolved cases are moved to the unscored set. AWS says fewer than 3% of scored samples remain contested after review, with a target of fewer than 1% surviving human review, and reports no labeling errors in a review of 100 randomly selected scored samples. These claims have not been independently confirmed here.

来源详情: helpnetsecurity.com ↗

为什么这很重要

The results suggest that strong performance on finding suspicious code does not necessarily translate into reliable vulnerability triage. High false-positive rates can burden security teams and weaken confidence in alerts, while stricter proof requirements can cause models to miss real vulnerabilities. The findings are relevant to organizations considering AI-assisted code review or security operations, but they do not measure complete commercial security products.

The addresses a practical weakness in AI-assisted security: a model may recognize a dangerous-looking pattern without understanding the controls that prevent exploitation. Help Net Security reports that direct prompting produced false-positive rates of 41% to 99%, while ranged from 52% to 71%.

Asking models to demonstrate that a vulnerability could actually be exploited reduced false positives by 17 to 74 percentage points, according to the report, but increased false-negative rates to between 7% and 44%. AWS’s stated minimum production bar was below 10% for both measures, and none of the tested configurations met both thresholds.

These results should not be read as a ranking of complete security products. AWS tested single-turn prompts on general-purpose models, not purpose-built systems that use tools, repeated validation, or agentic workflows. The source therefore supports caution about model-level vulnerability judgments, not a conclusion that AI security products as a whole fail.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下来看什么

Watch for independent evaluations using the released samples, results from tool-using or multi-step security systems, and evidence about performance in real production environments. The source does not document pricing, service-level access, or whether AWS’s scoring process is available without restrictions. It also does not establish that the tested models perform similarly on undisclosed enterprise code.

Independent researchers can use the public samples and evaluation process without recreating AWS’s reported data-generation costs, but withheld labels and AWS-verified scoring are intended to limit -specific optimization. Replication by outside groups would help test the reported results and labeling quality.

Future evaluations should compare single-pass models with systems that inspect repositories, execute code safely, reason over deployment configuration, and require evidence before issuing an alert. Those tests would better indicate how much practical security tooling improves on the model-only results.

The source leaves important unknowns: it does not identify the 12 tested model configurations, report per-model results in the supplied text, document pricing or access conditions for verified scoring, or show how performance transfers to proprietary codebases and live security operations.

相关指南和测验

人工智能模型解释AI 伦理Prompt Engineering测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注AI监管追踪器
觉得这有用吗?