Subira ku makuru
UmutekanoAI Understanding ibisobanuro

Ibipimo bya AWS birasanga ibizamini bya AI byerekana intege nke

Fasha Net Security ivuga ko AWS yagerageje moderi 12 za AI ku gipimo cyagenewe gutandukanya intege nke zishobora gukoreshwa na code itekanye gusa iteje akaga. Ntanumwe wujuje AWS imbibi zerekana umusaruro kubinyoma-byiza nibinyoma-bibi.

4 min readRead the linked source
Source-provided image accompanying AWS benchmark finds AI vulnerability detectors flag safe code
InkomokoInkomoko yanditse
Umwanditsi
helpnetsecurity.com
Ihuza ry'inkomoko
helpnetsecurity.comhttps://www.helpnetsecurity.com/2026/09/14/aws-deception-benchmark-security-vulnerabilities/
Ubwoko bw'inkomoko
Inkomoko ihujwe - ibanze-isoko yimiterere ntabwo yashizweho.
ImirongoSobanukirwa ibi mumasegonda 60

Tangira hano

Amagambo y'ingenzi

Ibipimo
Ikizamini gisanzwe cyangwa dataset ikoreshwa mugupima no kugereranya imikorere yicyitegererezo.
Icyitonderwa
Umubare wibyiza byahanuwe mubyukuri nibyo.
IsuzumeModeri ya AI Yasobanuwe Ikibazo

Byagenze bite

Help Net Security reports that AWS has publicly released its Deception , which tests whether AI models can distinguish genuine software vulnerabilities from safe code containing misleading vulnerability patterns. The benchmark includes 14,822 samples spanning 16 programming languages and more than 70 CWE categories; 9,695 are scored and 5,127 are mixed in as unscored samples.

Help Net Security reports that AWS evaluated 12 general-purpose models from five providers using a designed around false positives. The benchmark’s safe examples contain real vulnerability patterns alongside protections that prevent exploitation. Some challenges vary deployment conditions, such as Kubernetes network policies, so a model must consider both the code and its environment.

According to the report, AWS generated and refined examples against frontier models, excluding samples that were too easy to classify. AWS says this process consumed tens of billions of tokens. The company publicly releases the samples but withholds their labels; users submit predictions to AWS for verified scoring. The source does not state whether access to scoring carries a fee or other eligibility requirements.

Help Net Security reports that independent reviewers check labels without seeing one another’s decisions or the original reasoning. Disputed samples receive further review, and unresolved cases are moved to the unscored set. AWS says fewer than 3% of scored samples remain contested after review, with a target of fewer than 1% surviving human review, and reports no labeling errors in a review of 100 randomly selected scored samples. These claims have not been independently confirmed here.

Ibisobanuro birambuye: helpnetsecurity.com ↗

Impamvu ari ngombwa

The results suggest that strong performance on finding suspicious code does not necessarily translate into reliable vulnerability triage. High false-positive rates can burden security teams and weaken confidence in alerts, while stricter proof requirements can cause models to miss real vulnerabilities. The findings are relevant to organizations considering AI-assisted code review or security operations, but they do not measure complete commercial security products.

The addresses a practical weakness in AI-assisted security: a model may recognize a dangerous-looking pattern without understanding the controls that prevent exploitation. Help Net Security reports that direct prompting produced false-positive rates of 41% to 99%, while ranged from 52% to 71%.

Asking models to demonstrate that a vulnerability could actually be exploited reduced false positives by 17 to 74 percentage points, according to the report, but increased false-negative rates to between 7% and 44%. AWS’s stated minimum production bar was below 10% for both measures, and none of the tested configurations met both thresholds.

These results should not be read as a ranking of complete security products. AWS tested single-turn prompts on general-purpose models, not purpose-built systems that use tools, repeated validation, or agentic workflows. The source therefore supports caution about model-level vulnerability judgments, not a conclusion that AI security products as a whole fail.

Interactive Mechanism

Uburyo bukoreshwa: Uburyo bukora

Shakisha ikoranabuhanga ryihishe inyuma yiri terambere.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Kugenzura Ibitekerezo Byagenzuwe+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Ibyo kureba

Watch for independent evaluations using the released samples, results from tool-using or multi-step security systems, and evidence about performance in real production environments. The source does not document pricing, service-level access, or whether AWS’s scoring process is available without restrictions. It also does not establish that the tested models perform similarly on undisclosed enterprise code.

Independent researchers can use the public samples and evaluation process without recreating AWS’s reported data-generation costs, but withheld labels and AWS-verified scoring are intended to limit -specific optimization. Replication by outside groups would help test the reported results and labeling quality.

Future evaluations should compare single-pass models with systems that inspect repositories, execute code safely, reason over deployment configuration, and require evidence before issuing an alert. Those tests would better indicate how much practical security tooling improves on the model-only results.

The source leaves important unknowns: it does not identify the 12 tested model configurations, report per-model results in the supplied text, document pricing or access conditions for verified scoring, or show how performance transfers to proprietary codebases and live security operations.

Ibijyanye nuyobora & ibibazo

Moderi ya AI YasobanuweImyitwarire ya AIPrompt EngineeringGerageza ibyo uzi - gerageza ikibazo cya AI kubuntuReba ijambo AI mumagambo yacuKurikiza inzira ya AI ikurikirana
Basanze ari ingirakamaro?