NVIDIA Vera Rubin Platform Enters Production With CoreWeave DeepSeek R1 Benchmarks And AI Efficiency

NVIDIA Vera Rubin Platform Enters Production With CoreWeave DeepSeek R1 Benchmarks And AI Efficiency

NVIDIA has officially launched the Vera Rubin platform into full production. The new rack scale system pairs custom processing with high speed networking to target massive artificial intelligence installations. Early benchmarks suggest the architecture delivers efficiency gains over older chip setups, with production ramping up immediately at partner facilities like CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure.

Early test results show a massive jump in capability. Running a DeepSeek R1 benchmark on live hardware, CoreWeave recorded 10x more token throughput per megawatt compared to the older Grace Blackwell NVL72 system. Power is the primary limiting factor for modern data centers. This benchmark suggests operators can squeeze far more computational output from the same power grid.

This efficiency comes from tight integration. NVIDIA designed 7 chips and 5 distinct trays to work as 1 unified system. At the center sits the custom Vera CPU. The chip uses a custom Olympus core that provides 2x the single threaded performance and 3x the core to core bandwidth of older designs. It also drops memory latency by 40 percent. Cloud provider DeepInfra confirmed that the processor runs orchestration tasks 2x faster and supports 1.6x more concurrent AI agents.

Data must move fast to keep up with these processors. The platform uses generation 6 NVLink scale up technology to double throughput on complex workloads. For scale out networking, NVIDIA introduced the Spectrum 6 switch and the ConnectX 9 SuperNIC. Together they deliver 1.6x higher RDMA bandwidth than standard Ethernet options. Early adopters like SpaceXAI, Tesla, and Lambda are already installing these switches to run their operations.

Google Cloud is using the new platform for its fresh A5X bare metal instances. London based startup Ineffable Intelligence is the first to use these systems to train reinforcement learning models. Instead of learning from static databases, their software learns by interacting with simulated environments. This loop requires high memory bandwidth and low latency. The startup reports that the hardware was up and running almost immediately.

NVIDIA Vera Rubin Platform Enters Production With CoreWeave DeepSeek R1 Benchmarks And AI Efficiency

The hardware is also powering regional initiatives in Europe. Microsoft and Mistral announced a multibillion dollar partnership to build out European cloud systems. Under the deal, Mistral will add thousands of Vera Rubin GPUs to its data centers. This setup allows local governments and regulated industries to run advanced models under regional data laws. The transition is important because modern agent software consumes up to 15x more tokens than standard applications.

NVIDIA also updated the physical design to speed up assembly. The Vera Rubin NVL72 contains no cables, fans, or hoses inside the compute trays, dropping assembly time down to just 1 minute. The system uses a closed loop liquid cooling layout that tolerates a 45 degree Celsius water inlet. This allows operators to run data centers without expensive water chillers, saving millions of gallons of water per megawatt annually.

About the author

Majid T.
Owner of Technetbook | 10+ Years of Expertise in Technology | Seasoned Writer, Designer, and Programmer | Specialist in In-Depth Tech Reviews and Industry Insights | Passionate about Driving Innovation and Educating the Tech Community Technetbook

Join the conversation

Newsletter Subscription