뉴스로 돌아가기
제품AI Understanding 브리핑

Qualcomm, Always-On 에이전트 AI를 목표로 하는 Hexagon NPU 공개

Lec은 Qualcomm의 차세대 Hexagon NPU에 Element Accelerator, 50% 더 많은 공유 메모리 및 최대 300억 매개변수의 MoE 모델에 대한 최초의 지원을 추가했다고 보고했습니다.

4 min readRead the linked source
Source-provided image accompanying Qualcomm unveils Hexagon NPU aimed at always-on agentic AI
소스 참조녹음된 소스
출판사
thelec.net
소스 링크
thelec.nethttps://www.thelec.net/news/articleView.html?idxno=13884
소스 유형
연결된 소스 — 기본 소스 상태가 설정되지 않았습니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

전문가 혼합(MoE)
입력당 선택된 전문가만 실행하는 특수 하위 네트워크가 있는 아키텍처입니다.
메모리(에이전트 메모리)
AI 에이전트는 연속성을 향상하기 위해 여러 단계 또는 세션에서 사용하는 저장된 컨텍스트입니다.
변압기
시퀀스 전체의 관계를 병렬로 모델링하는 데 주의를 기울이는 신경 아키텍처입니다.
자신을 테스트해 보세요AI 에이전트 퀴즈

무슨 일이 일어났나요?

The Lec reports that Qualcomm unveiled a next-generation Hexagon neural processing unit on September 10 for continuous agentic AI workloads on mobile devices. The reported design adds an Element Accelerator for workloads, increases shared memory by 50%, and supports mixture-of-experts models with up to 30 billion parameters, with roughly 3 billion active for each generated token. According to The Lec, the NPU supports FP16, INT2, INT4, INT8 and FP8 data formats. Qualcomm claims INT4 prefill performance can be up to 50% faster and that responses can begin within 1.5 seconds, although The Lec does not report independent testing of those figures. The NPU is intended to work with Snapdragon’s Oryon CPU and Adreno GPU for multistep workflows and graphics-related AI processing. The report says Qualcomm will provide more details at Snapdragon Summit in the United States from September 22 to 24. The source does not identify a specific smartphone, launch date, retail availability, price or confirmed third-party benchmark results.

The Lec reports that Qualcomm unveiled its next-generation Hexagon NPU on September 10, positioning it as mobile hardware for AI systems that interpret context, reason across applications and perform tasks for users. The reported hardware adds an Element Accelerator alongside existing scalar, vector and matrix extensions, with the goal of improving -model computation while managing power use.

The report says shared memory capacity is 50% higher than in the prior design. Qualcomm’s stated rationale is that keeping model state, activations and intermediate tensor data closer to the processor can reduce transfers to external DRAM, potentially improving context retention and task switching. The report does not provide an independent comparison with an earlier Hexagon implementation.

The new Hexagon reportedly supports mixture-of-experts models with up to 30 billion parameters, while activating about 3 billion routed parameters per token. The Lec also reports NAND flash-to-memory management and caching intended to load only the relevant expert components into DRAM and retain frequently used components in cache.

The NPU reportedly supports FP16, INT2, INT4, INT8 and FP8 precision formats. The Lec attributes claims of up to 50% faster INT4 prefill and response starts within 1.5 seconds to Qualcomm. No independent tests, device availability, pricing or specific consumer products are provided.

소스 세부정보: thelec.net ↗

왜 중요한가요?

If Qualcomm’s reported design reaches consumer devices as described, it could make larger and more context-aware AI systems more practical to run locally on phones. Keeping more model data near the processor may reduce memory traffic and latency, while mixture-of-experts architectures could provide access to specialized capabilities without activating every parameter for every task. That matters for privacy, responsiveness and offline use, but the practical benefit remains unverified until devices, models and independent tests are available.

The announcement targets a central constraint in mobile AI: running capable models within tight limits on memory bandwidth, battery use and latency. A larger shared-memory pool and selective activation of mixture-of-experts components could reduce the amount of data moved during local inference, if Qualcomm’s architecture performs as described.

Local processing can have practical benefits because some interactions may be handled on a phone without sending every input to a remote service. However, the source does not establish privacy guarantees, offline functionality, battery improvements or performance across real applications. Those outcomes depend on software, model compression, thermal limits and handset implementation.

The reported integration with the CPU and GPU suggests Qualcomm is treating agentic AI as a platform workload rather than an isolated accelerator feature. Its significance will depend on whether developers can access the hardware efficiently and whether phone makers ship models and applications that use it.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
대화형 개념 확인+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

다음에 무엇을 볼 것인가

The next meaningful evidence will be Qualcomm’s Snapdragon Summit disclosures and eventual devices using the NPU. Watch for named chip platforms, supported models, developer tools, actual on-device measurements, sustained power consumption and whether the claimed response times apply to complete agentic workflows rather than limited demonstrations. Availability, pricing and regional distribution are not documented in the report.

Qualcomm says further details will be revealed at Snapdragon Summit from September 22 to 24. Those disclosures may clarify the specific Snapdragon platform, deployment timeline, supported operating environments and developer access.

Independent testing should verify the reported INT4 performance improvement, 1.5-second response-start claim, power consumption and sustained performance under multistep agentic workloads. The source provides no such testing.

It is not yet known which smartphones will use the NPU, when they will reach customers, what they will cost or whether all reported capabilities will be enabled at launch.

관련 가이드 및 퀴즈

AI 에이전트AI 모델 설명트랜스포머알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.AI 모델 출시 추적기를 따르세요.
이것이 유용하다고 생각하시나요?