Voltar às notícias
ProdutoInstruções AI Understanding

NVIDIA says Groq 3 LPX enters full production for faster agentic-AI inference

NVIDIA says its Groq 3 LPX inference accelerator is now in full production and delivered 3,400 output tokens per second in a company-cited Artificial Analysis benchmark using Gemma 4 31B with a 100,000-token context. Nebius plans to be the first AI cloud to deploy it.

Por 6 min read
Primary-source image accompanying NVIDIA says Groq 3 LPX enters full production for faster agentic-AI inference
A versão curta

NVIDIA says its Groq 3 LPX inference accelerator is now in full production and delivered 3,400 output tokens per second in a company-cited Artificial Analysis benchmark using Gemma 4 31B with a 100,000-token context. Nebius plans to be the first AI cloud to deploy it.

O que aconteceu

NVIDIA announced that Groq 3 LPX, an inference accelerator designed for fast token generation, is now in full production as an extension of its Vera Rubin platform. The company cited a record result on one model and named Nebius as the first planned AI-cloud adopter.

NVIDIA said on August 24, 2026, that Groq 3 LPX is now in full production. The company describes it as an interactive AI inference accelerator and an extension of the Vera Rubin platform, specifically intended to increase the rate at which an individual system generates tokens. The announcement frames token generation as distinct from processing a large context: agentic systems need both to read substantial context and to produce responses quickly. NVIDIA says the accelerator is designed for the second task, extending Vera Rubin NVL72 systems with faster generation for context-heavy workloads.

The company reported that Groq 3 LPX produced 3,400 output tokens per second in an Artificial Analysis benchmark using Gemma 4 31B, described in the release as an open-source agentic model, with a 100,000-token context. NVIDIA called this the fastest performance ever recorded for that model. This is a company-reported description of a benchmark result cited in the release; the source does not provide the full test configuration, comparison methodology, run count, latency distribution or independent confirmation. It also does not establish that the result applies to other models or production workloads.

NVIDIA said Groq 3 LPX provides four-times-faster responsiveness than the nearest alternative platform for agents and latency-sensitive workloads. It also said agentic coding tasks could be completed in minutes rather than hours. Those are broad performance and use-case claims from NVIDIA, and the source does not specify the tasks, baselines, software settings or conditions behind them. The release explains the intended mechanism: faster token generation could give an agent more time to inspect files, write and test code, call tools, verify results and iterate during a fixed period.

Nebius is identified as the first AI cloud planning to adopt the accelerator. The company said it intends to bring Groq 3 LPX to Nebius Token Factory, its production inference platform, using the same API developers already use and without requiring a migration to a new stack. NVIDIA also said purpose-built AI inference cloud Groq plans to be among the earliest adopters after Nebius. The announcement does not provide a deployment date, pricing, customer availability, service-level commitments or technical details about how the accelerator will be integrated into either provider’s production service.

Leia a fonte primária: nvidianews.nvidia.com

Por que isso importa

The announcement targets a central bottleneck in agentic AI: the time required to generate tokens across many reasoning, tool-use and verification steps. If the reported performance transfers to broader workloads, it could make some latency-sensitive AI services more responsive, though the source does not establish independent results, pricing or general availability.

Agentic AI systems often perform work through repeated cycles rather than a single response. In the workflow NVIDIA describes, an agent may inspect information, generate an action, call a tool, evaluate the result and continue. Each cycle can depend on new token generation, so lower generation latency could reduce the time users wait for multi-step tasks. The source’s practical claim is therefore narrower than a general claim that an AI model is more capable: Groq 3 LPX is intended to make inference more responsive when token generation is the limiting factor.

The announcement also illustrates a division of labor within AI infrastructure. NVIDIA presents Vera Rubin NVL72 as a broad training and inference platform, while Groq 3 LPX is positioned as a specialized extension for interactive generation. The company says the wider platform combines multiple components, including Vera CPU racks, BlueField-4 DPUs, Vera BlueField-4 storage and Spectrum-6 Ethernet. That architecture suggests NVIDIA is selling a coordinated rack-scale system for different AI-factory workloads, not merely a standalone chip, although the source does not quantify the contribution or cost of each component.

For developers and enterprises, the potential benefit is responsiveness in applications where users or downstream tools must wait for each step. Faster generation could matter for coding agents, other tool-using systems and high-volume inference services, particularly when long contexts are involved. But speed alone does not demonstrate better answers, safer actions, lower total cost or higher task-completion rates. The source does not report accuracy, reliability, energy use, error rates, safety evaluations or results on real customer workloads.

The commercial significance depends on access. Nebius’s planned deployment could give developers a way to use the hardware through an existing API, which would reduce migration friction if the service becomes available as described. Still, “full production” for the accelerator does not by itself mean that any developer can immediately buy or access it. NVIDIA’s release says features, pricing, availability and specifications may change, and it includes forward-looking language covering expected performance, partner arrangements and product benefits. Those qualifications are important limits on what the announcement establishes today.

O que assistir a seguir

The key follow-up is whether Nebius makes Groq 3 LPX available through its Token Factory platform and whether independent tests reproduce NVIDIA’s result. Observers should also look for performance across different models, context sizes, workloads, rack configurations and power requirements.

First, verify the timing and scope of Nebius’s deployment. The release says Nebius plans to bring Groq 3 LPX to Token Factory and calls it the first AI cloud to do so, but it gives no launch date or availability terms. A meaningful update would identify whether access is public, restricted to selected customers or still in testing, and would specify the models, regions and service conditions supported.

Second, look for independent benchmarking. NVIDIA’s headline result uses Gemma 4 31B and a 100,000-token context, while its four-times claim refers to a nearest alternative platform without naming the comparison in the source. Reproducible tests should disclose batch size, concurrency, input and output lengths, latency measures, software versions, hardware configuration and whether performance is measured per accelerator, per rack or per user.

Third, assess the trade-off between speed, quality and operating cost. The release emphasizes output-token rate but does not report power consumption, total system cost, utilization or quality under the same workload. For production buyers, tokens per second may be less informative than completed tasks per dollar or per watt, especially if agents spend substantial time on tool calls, verification or external services.

Finally, watch whether the accelerator’s benefits extend beyond the cited model and coding scenario. The source describes agentic systems broadly, but the evidence presented is a single company-cited benchmark result and product claims about latency-sensitive workloads. Future evidence should cover multiple model sizes, context lengths, concurrent users and long-running workflows, while also reporting failures and reliability. Until then, the announcement supports a significant infrastructure and production milestone, not a general conclusion about agent capability.

Guias e questionários relacionados

Agentes de IAModelos de IA explicadosTreinamento de IAFuturo da IATeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?