Назад до новин
БезпекаAI Understanding брифінг

Новий контрольний тест показує агентські структури ШІ, вразливі до мультимодального швидкого впровадження

Дослідники представляють MMPIBench, відтворюваний еталонний тест, який показує, що в той час як агентські структури штучного інтелекту часто блокують візуальні підказки на етапі планування, атаки на основі аудіо вдаються майже в половині протестованих сценаріїв, де підтримується канал.

4 min readRead the primary source
Source-provided image accompanying New benchmark reveals agentic AI frameworks vulnerable to multimodal prompt injection
Першоджерельний документДжерело записано
Видавець
arxiv.org
Посилання на джерело
arxiv.orghttps://arxiv.org/abs/2609.09404
Тип джерела
Первинний документ — офіційне оголошення, папір, документ або сторінка першої сторони, яку ми безпосередньо читаємо.
КонтекстЗрозумійте це за 60 секунд

Почніть тут

Ключові терміни

Оперативна ін'єкція
Шаблон атаки, коли зловмисні інструкції вставляються у вхідні дані моделі або отриманий вміст.
Еталон
Стандартизований тест або набір даних, який використовується для вимірювання та порівняння продуктивності моделі.
Підкажіть
Вхідні інструкції та контекст, надані для генеративної моделі.
Перевір себеВікторина агентів ШІ

Що сталося

Researchers published MMPIBench, a evaluating multimodal attacks on agentic AI frameworks. The study tested six frameworks and five foundation models across visual and audio channels, finding that while visual attacks are largely blocked during planning, audio attacks have a significantly higher completion rate when the infrastructure supports them.

Researchers introduced MMPIBench, a reproducible designed to measure the impact of multimodal attacks on agentic AI frameworks. These frameworks allow language models to plan, maintain memory, and call tools that interact with real-world files, emails, and services. The benchmark delivers a fixed set of attacks through six visual carriers, including OCR text, overlays, EXIF metadata, QR codes, fake interfaces, and hybrids, tracking how far injected instructions travel from perception through planning to tool execution.

The study conducted 720 runs covering six frameworks, five foundation models, six carriers, and four attacker objectives. Results showed that attacks completed in approximately 1% of runs but were attempted in 12.8%. The gap between attempted and completed attacks was closed almost entirely at the planning step, where the model read the injected instruction and declined to act. The specific foundation model used mattered far more than the framework for whether an instruction was acted upon; one model never attempted an attack and recognized the injection in 59.7% of runs, while two others attempted attacks in 23.6% of runs.

The researchers extended the to audio, the only other raw perceptual channel accepted by current frontier models. Only two of the five models ingested audio, and only three of the six frameworks delivered it. However, where the signal arrived, the attack completed in 49% of cells, rising to 75% for one specific model. This indicates that while visual channels are heavily defended by planning logic, audio channels are narrower but much less defended, leading to higher success rates for malicious instructions.

Деталі джерела: arxiv.org ↗

Чому це важливо

This research provides concrete, reproducible evidence that agentic AI systems, which can access real files and services, remain vulnerable to indirect via non-textual channels. The findings highlight a critical security gap: while visual injections are often neutralized by model reasoning, audio channels are less defended and can lead to successful malicious tool calls. This underscores the need for robust input sanitization and multi-modal safety training in agentic deployments.

The study demonstrates that reporting only attack completion rates understates the actual security exposure of agentic AI systems. The high rate of attempted attacks (12.8%) compared to completed ones (1%) suggests that current models often recognize malicious intent but that this defense is not uniform across all models or modalities.

The finding that audio-based attacks have a 49% completion rate in supported environments is particularly concerning because audio is a less common input channel for many agentic deployments, meaning it may receive less security scrutiny. This creates a potential blind spot where attackers could bypass visual defenses by using audio inputs to manipulate agents into executing harmful tool calls.

The variability between models, with one model recognizing injections in nearly 60% of runs and others attempting attacks in over 20%, highlights that model selection is a critical factor in agentic AI security. Organizations cannot rely solely on framework-level safeguards and must consider the specific safety characteristics of the underlying foundation models.

Interactive Mechanism

Інтерактивний механізм: як він насправді працює

Дослідіть технологію, що лежить в основі цієї розробки, в інтерактивному режимі.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Інтерактивна перевірка концепції+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Що дивитися далі

Developers and enterprises deploying agentic AI should monitor for updates to MMPIBench and related security patches. The industry may see increased focus on audio input sanitization and stricter permissioning for tool calls in response to these findings.

Monitor for updates to MMPIBench and subsequent research that may test additional perceptual channels or more sophisticated attack vectors. The reproducibility of the allows for ongoing community verification and expansion.

Watch for industry responses from major AI framework developers and foundation model providers, who may release security patches or updated safety training to address the specific vulnerabilities identified in the audio and visual channels.

Observe how enterprises deploying agentic AI adjust their security protocols, potentially implementing stricter input filtering for non-textual data and more granular permission controls for tool calls to mitigate the risks highlighted by the study.

Пов’язані посібники та вікторини

Агенти ШІЕтика ШІПояснення моделей AIПеревірте свої знання — пройдіть безкоштовну вікторину зі штучним інтелектомЗнайдіть термін ШІ в нашому глосаріїДотримуйтесь трекера регулювання ШІ
Знайшли це корисним?