Zpět na Novinky
ZabezpečeníInstruktáž AI Understanding

Benchmark AWS najde detektory zranitelnosti umělé inteligence bezpečný kód

Help Net Security hlásí, že AWS testoval 12 modelů umělé inteligence v benchmarku navrženém k rozlišení zneužitelných zranitelností od bezpečného kódu, který jen nebezpečně vypadá. Žádná nesplňovala stanovený práh produkce AWS pro falešně pozitivní i falešně negativní míru.

4 min readRead the linked source
Source-provided image accompanying AWS benchmark finds AI vulnerability detectors flag safe code
Odkaz na zdrojZdroj zaznamenán
Vydavatel
helpnetsecurity.com
Odkaz na zdroj
helpnetsecurity.comhttps://www.helpnetsecurity.com/2026/09/14/aws-deception-benchmark-security-vulnerabilities/
Typ zdroje
Propojený zdroj — stav primárního zdroje nebyl stanoven.
KontextPochopte to za 60 sekund

Začněte zde

Klíčové pojmy

Benchmark
Standardizovaný test nebo soubor dat používaný k měření a porovnávání výkonu modelu.
Přesnost
Podíl předpokládaných pozitiv, které jsou skutečně správné.
Otestujte seKvíz s vysvětlením modelů umělé inteligence

Co se stalo

Help Net Security reports that AWS has publicly released its Deception , which tests whether AI models can distinguish genuine software vulnerabilities from safe code containing misleading vulnerability patterns. The benchmark includes 14,822 samples spanning 16 programming languages and more than 70 CWE categories; 9,695 are scored and 5,127 are mixed in as unscored samples.

Help Net Security reports that AWS evaluated 12 general-purpose models from five providers using a designed around false positives. The benchmark’s safe examples contain real vulnerability patterns alongside protections that prevent exploitation. Some challenges vary deployment conditions, such as Kubernetes network policies, so a model must consider both the code and its environment.

According to the report, AWS generated and refined examples against frontier models, excluding samples that were too easy to classify. AWS says this process consumed tens of billions of tokens. The company publicly releases the samples but withholds their labels; users submit predictions to AWS for verified scoring. The source does not state whether access to scoring carries a fee or other eligibility requirements.

Help Net Security reports that independent reviewers check labels without seeing one another’s decisions or the original reasoning. Disputed samples receive further review, and unresolved cases are moved to the unscored set. AWS says fewer than 3% of scored samples remain contested after review, with a target of fewer than 1% surviving human review, and reports no labeling errors in a review of 100 randomly selected scored samples. These claims have not been independently confirmed here.

Podrobnosti o zdroji: helpnetsecurity.com ↗

Proč na tom záleží

The results suggest that strong performance on finding suspicious code does not necessarily translate into reliable vulnerability triage. High false-positive rates can burden security teams and weaken confidence in alerts, while stricter proof requirements can cause models to miss real vulnerabilities. The findings are relevant to organizations considering AI-assisted code review or security operations, but they do not measure complete commercial security products.

The addresses a practical weakness in AI-assisted security: a model may recognize a dangerous-looking pattern without understanding the controls that prevent exploitation. Help Net Security reports that direct prompting produced false-positive rates of 41% to 99%, while ranged from 52% to 71%.

Asking models to demonstrate that a vulnerability could actually be exploited reduced false positives by 17 to 74 percentage points, according to the report, but increased false-negative rates to between 7% and 44%. AWS’s stated minimum production bar was below 10% for both measures, and none of the tested configurations met both thresholds.

These results should not be read as a ranking of complete security products. AWS tested single-turn prompts on general-purpose models, not purpose-built systems that use tools, repeated validation, or agentic workflows. The source therefore supports caution about model-level vulnerability judgments, not a conclusion that AI security products as a whole fail.

Interactive Mechanism

Interaktivní mechanismus: Jak to vlastně funguje

Interaktivně prozkoumejte základní technologii tohoto vývoje.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interaktivní kontrola konceptu+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Na co se dále dívat

Watch for independent evaluations using the released samples, results from tool-using or multi-step security systems, and evidence about performance in real production environments. The source does not document pricing, service-level access, or whether AWS’s scoring process is available without restrictions. It also does not establish that the tested models perform similarly on undisclosed enterprise code.

Independent researchers can use the public samples and evaluation process without recreating AWS’s reported data-generation costs, but withheld labels and AWS-verified scoring are intended to limit -specific optimization. Replication by outside groups would help test the reported results and labeling quality.

Future evaluations should compare single-pass models with systems that inspect repositories, execute code safely, reason over deployment configuration, and require evidence before issuing an alert. Those tests would better indicate how much practical security tooling improves on the model-only results.

The source leaves important unknowns: it does not identify the 12 tested model configurations, report per-model results in the supplied text, document pricing or access conditions for verified scoring, or show how performance transfers to proprietary codebases and live security operations.

Související průvodci a kvízy

Vysvětlení modelů AIEtika AIPrompt EngineeringOtestujte si, co víte – vyzkoušejte bezplatný kvíz AIVyhledejte si termín AI v našem slovníkuSledujte sledovač regulace AI
Považujete to za užitečné?