뉴스로 돌아가기
보안AI Understanding 브리핑

조선비즈(ChosunBiz)는 Microsoft 보안 임원이 멀티 에이전트 AI가 취약성 탐지를 향상한다고 보도했습니다.

조선비즈에 따르면 Microsoft 김태수 부사장은 멀티 에이전트 워크플로우가 단일 AI 시스템보다 소프트웨어 취약점을 더 효율적으로 찾아 검증할 수 있다고 말했다. 이 보고서는 DARPA의 AI Cyber Challenge와 Microsoft의 MDASH 시스템의 결과를 설명하지만 주장은 독립적이지 않습니다.

6 min readRead the original reporting
Primary-source image accompanying ChosunBiz reports Microsoft security executive says multi-agent AI improves vulnerability detection
기여 보고녹음된 소스
출판사
biz.chosun.com
소스 링크
biz.chosun.comhttps://biz.chosun.com/en/en-it/2026/08/26/DNA3TDFF75HF5IFJD7IL6HJ6WE/?outputType=amp
소스 유형
자사 문서가 아닌 뉴스 매체를 통한 보도입니다.

자체적으로는 확인할 수 없었던 내용: 이 소유권 주장은 해당 매장에 귀속됩니다. 당사는 자사 문서와 비교하여 이를 확인하지 않았습니다. (biz.chosun.com)

맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
컴퓨팅
모델을 훈련하고 실행하는 데 필요한 처리 리소스는 FLOPS 또는 GPU 시간으로 측정되는 경우가 많습니다.
대기 시간
요청을 보내는 것과 모델의 출력을 받는 것 사이의 시간입니다.
자신을 테스트해 보세요AI 에이전트 퀴즈

무슨 일이 일어났나요?

According to ChosunBiz, Kim Tae-su said specialized AI agents working together can improve vulnerability discovery, verification, exploit testing, and patch generation. The report cites Team Atlanta’s performance in DARPA’s AI Cyber Challenge and describes Microsoft’s MDASH system, which uses more than 100 sub-agents to analyze software.

ChosunBiz reported that Kim Tae-su, Microsoft’s corporate vice president for security and a Georgia Institute of Technology computer-engineering professor, made the comments during a keynote at SMARTCLOUD SHOW 2026 in Seoul. The report said his lecture covered “hyperscale bug finding,” including work connected to DARPA’s AI Cyber Challenge and Microsoft’s MDASH vulnerability-detection system. ChosunBiz presented Kim’s remarks as the basis for the article; the source does not provide an independent transcript, technical paper, or outside assessment of the systems described.

The report said Kim led Team Atlanta, the winner of last year’s AI Cyber Challenge, a DARPA competition focused on using AI to analyze, detect, and fix software vulnerabilities. ChosunBiz reported that Team Atlanta scored 392.76 points in the final, more than 170 points ahead of the runner-up. In the account, the team’s cyber reasoning system analyzed 28 open-source projects, found 54 of 70 vulnerabilities placed by organizers, fixed most vulnerabilities within an hour, and discovered 18 previously unknown zero-day vulnerabilities. ChosunBiz also reported that more than 90% of the discovered vulnerabilities confirmed as real, although the article does not explain the evaluation methodology in detail.

ChosunBiz reported that the system cost about $152 per successful task, citing Kim, who compared that figure with the much higher cost of human code audits or penetration tests. The report said Team Atlanta included researchers from Georgia Tech, KAIST, POSTECH, and other institutions, with 80% to 90% of the team described as Korean. These figures and comparisons are reported claims from Kim and ChosunBiz, not independently verified measurements in the supplied material.

The report said Kim described limitations in using one AI system to analyze very large codebases. In response, the team used a multi-agent architecture assigning different roles to specialized agents, including vulnerability hunting, static analysis, verification, and patch generation. ChosunBiz said Microsoft’s MDASH extends this idea by first building a threat model from software architecture and past vulnerabilities, then deploying more than 100 sub-agents. Agents acting as hackers, developers, and defenders reportedly check findings against one another, generate proof-of-concept code to test exploitability, and produce remediation patches.

ChosunBiz reported that Kim said MDASH had identified high-risk issues that could potentially allow an external attacker to obtain the highest system privileges by sending packets. Kim also said MDASH performed comparably to newer frontier models despite combining models two generations older, and that it was being used beyond benchmarks in real software-development environments. The supplied report does not identify those environments, disclose the models or test sets involved, provide examples of the vulnerabilities, or include independent confirmation of the claimed performance.

소스 세부정보: biz.chosun.com ↗

왜 중요한가요?

If independently validated, the approach could lower the cost of security audits and help defenders examine software more quickly. It could also increase the volume of vulnerabilities discovered, creating pressure for software maintainers to verify, prioritize, and fix findings faster.

The practical significance is the possibility of moving parts of vulnerability research from a labor-intensive, episodic process toward continuous automated analysis. ChosunBiz reported Kim’s estimate that AI-based work could reduce the expense of successful security tasks dramatically compared with human audits and penetration tests. Lower costs could make deeper review more accessible to organizations that cannot routinely hire specialist security teams, although the source does not establish that the reported price comparison applies across different software types or operating conditions.

A multi-agent design could address a known operational problem: a general-purpose model may reason well on one task but perform inconsistently on another. In ChosunBiz’s account, separate agents are assigned narrower responsibilities and then asked to verify one another’s findings. That structure may make errors easier to detect, but it also adds coordination, monitoring, and evaluation requirements. The source does not provide enough information to determine whether the additional agents improve accuracy consistently or simply increase and operational complexity.

The defensive value could be substantial if systems reliably identify vulnerabilities before attackers do. Faster discovery and patch generation could reduce the time that flaws remain exploitable, particularly in widely used open-source components. At the same time, the same capabilities may be useful to attackers. The article’s description of proof-of-concept generation and packet-triggered privilege-escalation findings illustrates why organizations would need strict controls over access, testing environments, disclosure, and the handling of exploit details.

The report also points to a broader shift in cybersecurity evaluation. Kim said AI models had surpassed human hackers in a March analysis, according to ChosunBiz, and argued that better AI could both make software safer and reveal more weaknesses. That is an important distinction: finding more vulnerabilities does not by itself mean security has improved. The public impact depends on whether maintainers can reproduce findings, prioritize genuine risks, deploy safe patches, and prevent newly generated exploit knowledge from being misused.

The evidence remains limited in the supplied source. ChosunBiz is a reported-secondary outlet, and the article itself says it was translated by AI. No primary MDASH documentation, competition data release, code, protocol, or independent researcher response is included. The reported results therefore merit attention as an account of Microsoft’s and Kim’s claims, but they should not yet be treated as a general proof that multi-agent systems outperform human security teams or single-model systems in real-world conditions.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
대화형 개념 확인+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

다음에 무엇을 볼 것인가

The key questions are whether MDASH’s results hold outside benchmarks, how often it misses real vulnerabilities or produces false positives, and how safely its proof-of-concept and patch-generation capabilities are controlled. Public technical documentation and independent testing would clarify the claims.

The first priority is independent replication. Security researchers and software maintainers would need access to clear evaluation procedures, representative codebases, vulnerability labels, and definitions of successful detection, confirmation, and repair. Particular attention should go to the reported 54-of-70 discovery result, the 18 zero-day findings, the more-than-90% confirmation rate, and the $152-per-task estimate. Without methodological detail, those numbers cannot be compared reliably with human audits, automated scanners, or other AI systems.

The next issue is error quality rather than raw discovery volume. Organizations should examine false positives, missed vulnerabilities, duplicate findings, incorrect severity ratings, and patches that compile but change intended behavior or create new flaws. ChosunBiz reported that agents with different perspectives reach consensus, but the source does not show how consensus is measured or whether it correlates with real-world correctness. Independent testing should assess both secure patching and the consequences of an incorrect automated fix.

Deployment safeguards will also matter. A system that generates proof-of-concept exploit code needs isolation, authorization controls, audit logs, and clear rules for responsible disclosure. The source reports that MDASH is being used in real software-development environments, but does not specify the environments, user permissions, human review procedures, or whether generated patches are automatically applied. Those details would determine whether the system reduces risk or introduces new pathways for misuse.

Researchers should compare the multi-agent approach with simpler alternatives, including a single strong model, conventional static-analysis tools, human-led penetration testing, and combinations of these methods. ChosunBiz reported that MDASH achieved performance comparable to newer frontier models while using older models, but did not provide the design or explain whether cost, , coverage, or reliability was the comparison’s main measure. Those distinctions are important for procurement and security planning.

Finally, maintainers and policymakers should watch whether AI-assisted vulnerability discovery changes disclosure workloads. More findings could improve software security if they are accurate and responsibly reported, but could also overwhelm open-source projects and smaller vendors. Meaningful unknowns include the system’s coverage across programming languages and application types, its performance on previously unseen code, the frequency of dangerous patch errors, and the extent of human oversight required before any result is acted upon.

관련 가이드 및 퀴즈

AI 에이전트AI 모델 설명AI 윤리AI 트레이닝알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 규제 추적기를 따르세요
이것이 유용하다고 생각하시나요?