NVIDIA has launched the Vera Rubin NVL72, a rack scale supercomputer designed to power the next generation of AI factories. This massive system integrates 72 Rubin GPUs, 36 Vera CPUs, ConnectX 9 SuperNICs, and BlueField 4 DPUs into a single unified architecture. But full production systems are currently ramping through server manufacturers, meaning immediate deployment is limited to select cloud hyperscalers.
According to details from the official NVIDIA product launch sheets, this architecture relies on the 3rd generation MGX NVL72 design to simplify upgrades from older platforms. The full rack configuration boasts 3168 custom NVIDIA Olympus cores, which are compatible with Arm instruction sets. Memory performance reaches new heights with 20.7 TB of HBM4 GPU memory pushing 1580 TB/s of bandwidth, alongside 54 TB of LPDDR5X CPU system memory. Scale out networking is handled by Quantum X800 InfiniBand and Spectrum X Ethernet, delivering 28.8 TB/s of aggregate bandwidth.
NVIDIA is throwing down the gauntlet with these efficiency numbers. The new system trains mixture of experts models with just 25% of the GPUs required by the previous GB200 Blackwell system. Training is much faster. It also cuts inference costs down to just 10% of the cost per million tokens compared to Blackwell. By boosting token throughput up to 10x per megawatt, the rack design squeezes massive performance out of the same physical power footprint.
Agentic AI workloads consume immense amounts of data. To solve this, the NVL72 can be paired with Groq 3 LPX racks. This combination delivers up to 35x higher throughput per megawatt for trillion parameter models. The setup is designed for real time token processing, handling million token contexts with ease. The processing pipeline also manages scientific computing with a smaller NVL4 configuration, which packs 4 Rubin GPUs and 2 Vera CPUs to outpace older Grace Hopper chips by up to 8x in deep learning inference.
Server makers in Taiwan are already starting mass assembly. Top tier supply chain partners are manufacturing the 7 new chips that make up the platform. The hardware includes the standalone Rubin GPU, which delivers 50 PFLOPS of NVFP4 inference, and the Vera Rubin Superchip, which combines 2 Rubin GPUs and 1 Vera CPU on a single board. Cloud providers and large laboratories will begin receiving these systems as production scales up over the coming months.
