Torna alle notizie
SicurezzaAI Understanding briefing

Il benchmark AWS rileva che i rilevatori di vulnerabilità AI contrassegnano il codice sicuro

Help Net Security segnala che AWS ha testato 12 modelli di intelligenza artificiale su un benchmark progettato per distinguere le vulnerabilità sfruttabili dal codice sicuro che sembra solo pericoloso. Nessuno ha soddisfatto la soglia di produzione dichiarata da AWS sia per il tasso di falsi positivi che per quello di falsi negativi.

4 min readRead the linked source
Source-provided image accompanying AWS benchmark finds AI vulnerability detectors flag safe code
Riferimento alla fonteFonte registrata
Editore
helpnetsecurity.com
Collegamento alla fonte
helpnetsecurity.comhttps://www.helpnetsecurity.com/2026/09/14/aws-deception-benchmark-security-vulnerabilities/
Tipo di fonte
Fonte collegata: lo stato di fonte primaria non è stato stabilito.
ContestoComprendilo in 60 secondi

Inizia qui

Termini chiave

Punto di riferimento
Un test o un set di dati standardizzato utilizzato per misurare e confrontare le prestazioni del modello.
Precisione
La percentuale di positivi previsti che sono effettivamente corretti.
Mettiti alla provaQuiz sulla spiegazione dei modelli di intelligenza artificiale

Cosa è successo

Help Net Security reports that AWS has publicly released its Deception , which tests whether AI models can distinguish genuine software vulnerabilities from safe code containing misleading vulnerability patterns. The benchmark includes 14,822 samples spanning 16 programming languages and more than 70 CWE categories; 9,695 are scored and 5,127 are mixed in as unscored samples.

Help Net Security reports that AWS evaluated 12 general-purpose models from five providers using a designed around false positives. The benchmark’s safe examples contain real vulnerability patterns alongside protections that prevent exploitation. Some challenges vary deployment conditions, such as Kubernetes network policies, so a model must consider both the code and its environment.

According to the report, AWS generated and refined examples against frontier models, excluding samples that were too easy to classify. AWS says this process consumed tens of billions of tokens. The company publicly releases the samples but withholds their labels; users submit predictions to AWS for verified scoring. The source does not state whether access to scoring carries a fee or other eligibility requirements.

Help Net Security reports that independent reviewers check labels without seeing one another’s decisions or the original reasoning. Disputed samples receive further review, and unresolved cases are moved to the unscored set. AWS says fewer than 3% of scored samples remain contested after review, with a target of fewer than 1% surviving human review, and reports no labeling errors in a review of 100 randomly selected scored samples. These claims have not been independently confirmed here.

Dettagli della fonte: helpnetsecurity.com ↗

Perché è importante

The results suggest that strong performance on finding suspicious code does not necessarily translate into reliable vulnerability triage. High false-positive rates can burden security teams and weaken confidence in alerts, while stricter proof requirements can cause models to miss real vulnerabilities. The findings are relevant to organizations considering AI-assisted code review or security operations, but they do not measure complete commercial security products.

The addresses a practical weakness in AI-assisted security: a model may recognize a dangerous-looking pattern without understanding the controls that prevent exploitation. Help Net Security reports that direct prompting produced false-positive rates of 41% to 99%, while ranged from 52% to 71%.

Asking models to demonstrate that a vulnerability could actually be exploited reduced false positives by 17 to 74 percentage points, according to the report, but increased false-negative rates to between 7% and 44%. AWS’s stated minimum production bar was below 10% for both measures, and none of the tested configurations met both thresholds.

These results should not be read as a ranking of complete security products. AWS tested single-turn prompts on general-purpose models, not purpose-built systems that use tools, repeated validation, or agentic workflows. The source therefore supports caution about model-level vulnerability judgments, not a conclusion that AI security products as a whole fail.

Interactive Mechanism

Meccanismo interattivo: come funziona realmente

Esplora la tecnologia alla base di questo sviluppo in modo interattivo.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Verifica concettuale interattiva+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Cosa guardare dopo

Watch for independent evaluations using the released samples, results from tool-using or multi-step security systems, and evidence about performance in real production environments. The source does not document pricing, service-level access, or whether AWS’s scoring process is available without restrictions. It also does not establish that the tested models perform similarly on undisclosed enterprise code.

Independent researchers can use the public samples and evaluation process without recreating AWS’s reported data-generation costs, but withheld labels and AWS-verified scoring are intended to limit -specific optimization. Replication by outside groups would help test the reported results and labeling quality.

Future evaluations should compare single-pass models with systems that inspect repositories, execute code safely, reason over deployment configuration, and require evidence before issuing an alert. Those tests would better indicate how much practical security tooling improves on the model-only results.

The source leaves important unknowns: it does not identify the 12 tested model configurations, report per-model results in the supplied text, document pricing or access conditions for verified scoring, or show how performance transfers to proprietary codebases and live security operations.

Guide e quiz correlati

Spiegazione dei modelli di intelligenza artificialeEtica dell'IAPrompt EngineeringMetti alla prova ciò che sai: prova un quiz gratuito sull'intelligenza artificialeCerca un termine AI nel nostro glossarioSegui il tracker della regolamentazione dell'IA
Lo hai trovato utile?