Επιστροφή στις Ειδήσεις
ΠροϊόνAI Understanding ενημέρωση

Η Qualcomm αποκαλύπτει το Hexagon NPU που στοχεύει στο πάντα ενεργό agent AI

Η Lec αναφέρει ότι η επόμενη γενιά Hexagon NPU της Qualcomm προσθέτει έναν επιταχυντή Element, 50% περισσότερη κοινή μνήμη και υποστήριξη για πρώτη φορά για μοντέλα MoE έως και 30 δισεκατομμυρίων παραμέτρων.

4 min readRead the linked source
Source-provided image accompanying Qualcomm unveils Hexagon NPU aimed at always-on agentic AI
Αναφορά πηγήςΗ πηγή καταγράφηκε
Εκδότης
thelec.net
Σύνδεσμος πηγής
thelec.nethttps://www.thelec.net/news/articleView.html?idxno=13884
Τύπος πηγής
Συνδεδεμένη πηγή — η κατάσταση της κύριας πηγής δεν έχει καθοριστεί.
ΠλαίσιοΚαταλάβετε αυτό σε 60 δευτερόλεπτα

Ξεκινήστε εδώ

Βασικοί όροι

Μίγμα Εμπειρογνωμόνων (ΥΠ)
Μια αρχιτεκτονική με εξειδικευμένα υποδίκτυα όπου εκτελούνται μόνο επιλεγμένοι ειδικοί ανά είσοδο.
Μνήμη (μνήμη πράκτορα)
Το αποθηκευμένο πλαίσιο που χρησιμοποιεί ένας πράκτορας τεχνητής νοημοσύνης σε βήματα ή περιόδους σύνδεσης για να βελτιώσει τη συνέχεια.
Μετασχηματιστής
Μια νευρωνική αρχιτεκτονική που χρησιμοποιεί την προσοχή για να μοντελοποιήσει τις σχέσεις μεταξύ των ακολουθιών παράλληλα.
Δοκιμάστε τον εαυτό σαςΚουίζ για πράκτορες AI

Τι έγινε

The Lec reports that Qualcomm unveiled a next-generation Hexagon neural processing unit on September 10 for continuous agentic AI workloads on mobile devices. The reported design adds an Element Accelerator for workloads, increases shared memory by 50%, and supports mixture-of-experts models with up to 30 billion parameters, with roughly 3 billion active for each generated token. According to The Lec, the NPU supports FP16, INT2, INT4, INT8 and FP8 data formats. Qualcomm claims INT4 prefill performance can be up to 50% faster and that responses can begin within 1.5 seconds, although The Lec does not report independent testing of those figures. The NPU is intended to work with Snapdragon’s Oryon CPU and Adreno GPU for multistep workflows and graphics-related AI processing. The report says Qualcomm will provide more details at Snapdragon Summit in the United States from September 22 to 24. The source does not identify a specific smartphone, launch date, retail availability, price or confirmed third-party benchmark results.

The Lec reports that Qualcomm unveiled its next-generation Hexagon NPU on September 10, positioning it as mobile hardware for AI systems that interpret context, reason across applications and perform tasks for users. The reported hardware adds an Element Accelerator alongside existing scalar, vector and matrix extensions, with the goal of improving -model computation while managing power use.

The report says shared memory capacity is 50% higher than in the prior design. Qualcomm’s stated rationale is that keeping model state, activations and intermediate tensor data closer to the processor can reduce transfers to external DRAM, potentially improving context retention and task switching. The report does not provide an independent comparison with an earlier Hexagon implementation.

The new Hexagon reportedly supports mixture-of-experts models with up to 30 billion parameters, while activating about 3 billion routed parameters per token. The Lec also reports NAND flash-to-memory management and caching intended to load only the relevant expert components into DRAM and retain frequently used components in cache.

The NPU reportedly supports FP16, INT2, INT4, INT8 and FP8 precision formats. The Lec attributes claims of up to 50% faster INT4 prefill and response starts within 1.5 seconds to Qualcomm. No independent tests, device availability, pricing or specific consumer products are provided.

Στοιχεία πηγής: thelec.net ↗

Γιατί έχει σημασία

If Qualcomm’s reported design reaches consumer devices as described, it could make larger and more context-aware AI systems more practical to run locally on phones. Keeping more model data near the processor may reduce memory traffic and latency, while mixture-of-experts architectures could provide access to specialized capabilities without activating every parameter for every task. That matters for privacy, responsiveness and offline use, but the practical benefit remains unverified until devices, models and independent tests are available.

The announcement targets a central constraint in mobile AI: running capable models within tight limits on memory bandwidth, battery use and latency. A larger shared-memory pool and selective activation of mixture-of-experts components could reduce the amount of data moved during local inference, if Qualcomm’s architecture performs as described.

Local processing can have practical benefits because some interactions may be handled on a phone without sending every input to a remote service. However, the source does not establish privacy guarantees, offline functionality, battery improvements or performance across real applications. Those outcomes depend on software, model compression, thermal limits and handset implementation.

The reported integration with the CPU and GPU suggests Qualcomm is treating agentic AI as a platform workload rather than an isolated accelerator feature. Its significance will depend on whether developers can access the hardware efficiently and whether phone makers ship models and applications that use it.

Interactive Mechanism

Διαδραστικός Μηχανισμός: Πώς λειτουργεί στην πραγματικότητα

Εξερευνήστε την υποκείμενη τεχνολογία πίσω από αυτήν την εξέλιξη διαδραστικά.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Διαδραστικός Έλεγχος Έννοιας+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Τι να παρακολουθήσετε στη συνέχεια

The next meaningful evidence will be Qualcomm’s Snapdragon Summit disclosures and eventual devices using the NPU. Watch for named chip platforms, supported models, developer tools, actual on-device measurements, sustained power consumption and whether the claimed response times apply to complete agentic workflows rather than limited demonstrations. Availability, pricing and regional distribution are not documented in the report.

Qualcomm says further details will be revealed at Snapdragon Summit from September 22 to 24. Those disclosures may clarify the specific Snapdragon platform, deployment timeline, supported operating environments and developer access.

Independent testing should verify the reported INT4 performance improvement, 1.5-second response-start claim, power consumption and sustained performance under multistep agentic workloads. The source provides no such testing.

It is not yet known which smartphones will use the NPU, when they will reach customers, what they will cost or whether all reported capabilities will be enabled at launch.

Σχετικοί οδηγοί και κουίζ

Πράκτορες AIΕπεξήγηση μοντέλων AIΜετασχηματιστέςΔοκιμάστε τι γνωρίζετε — δοκιμάστε ένα δωρεάν κουίζ AIΑναζητήστε έναν όρο AI στο γλωσσάρι μαςΑκολουθήστε τον ιχνηλάτη έκδοσης μοντέλου AI
Βρήκατε αυτό χρήσιμο;