
NVIDIA has moved its Groq 3 LPX inference accelerator into full production, targeting the massive computational demands of agentic AI. Built as an expansion for the Vera Rubin platform, the specialized hardware hits 3,400 output tokens per second in benchmark testing. Nebius is already signed on as the first cloud provider to deploy the silicon.
Agent systems rely on thousands of rapid reasoning steps, which makes raw token generation speeds critical for real time tasks. In benchmarks conducted by Artificial Analysis running the open source Gemma 4 31B model with a 100,000 token context, the Groq 3 LPX set a new performance ceiling. NVIDIA claims the accelerator provides 4x faster responsiveness compared to the closest competing systems, cutting long coding tasks down from hours to minutes.
The chip acts as an interactive companion to Vera Rubin NVL72 racks. While standard systems handle general processing, the LPX focuses purely on generation speeds so software agents can verify code and execute tools without delay. NVIDIA CEO Jensen Huang addressed the shift in workload demands during the Hot Chips presentation:
Vera Rubin extends that vision with workload optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This shifts how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide.
Nebius is integrating the hardware directly into its Token Factory platform, allowing engineers to access the accelerated speeds through existing software connections without restructuring their code. Purpose built provider Groq will also be among the earliest operators to integrate the units into its server fleet.
The broader hardware rollout spans 7 distinct chips across 5 dedicated rack configurations. These setups incorporate BlueField 4 data processing units, Vera CPU modules, Vera BlueField 4 STX storage arrays, and Spectrum 6 SPX ethernet networking to maintain low latency across multi agent environments.