On August 21, 2026, Microsoft confirmed that its data centers have received the first production NVIDIA Vera Rubin NVL72 racks, and NVIDIA announced that the platform is moving to full production. The Rubin generation is moving from lab validation to mass deliveries to hyperscalers, and the vendor-claimed figures — a fourfold lower GPU requirement for training 10-trillion-parameter MoE models and roughly 10x cheaper inference tokens — are getting a real production foundation for the first time.
What happened
On August 21, 2026, Microsoft CEO Satya Nadella posted photos of the delivery on X with the caption "Delivery day at our Microsoft DCs as the first production Vera Rubins arrive", confirming that the first production Vera Rubin NVL72 racks had arrived at the company's data centers. On the same day, NVIDIA's official account said the Vera Rubin platform is "ramping into full production". Microsoft/Azure became the first hyperscaler to receive production racks of the new generation.
Context
Vera Rubin NVL72 is NVIDIA's next rack-scale family of AI computers after Blackwell. A single rack contains 72 Rubin GPUs with HBM4 memory (50 PFLOPS in NVFP4 format), 36 Vera processors on Arm cores, an NVLink 6 switch with 3.6 TB/s bandwidth per GPU, ConnectX-9 network cards, and a BlueField-4 DPU; the platform combines six chips, including Spectrum-6 SPX. The rack is fully liquid-cooled, the trays are modular and cable-free, and installation time, according to Nvidia, has been reduced to 5 minutes versus two hours. Before this delivery, Vera Rubin existed in lab validation mode: in March 2026 at GTC, Microsoft first enabled NVL72 in its own lab. According to plans announced at GTC 2026, after Microsoft, Google Cloud and Lambda will follow in the second half of 2026, AWS, which will deploy more than 1 million GPUs, including Rubin, within 12 months, and neocloud clusters, including a 1.35 GW Nscale cluster for Microsoft.
Why this matters for the industry
The move of Vera Rubin from the lab to production deliveries sets an economic benchmark for the entire industry. According to Nvidia, compared with GB200 NVL72, training a 10-trillion-parameter MoE model requires four times fewer GPUs — 100 trillion tokens per month — and the cost of one million inference tokens for agentic tasks falls by about 10 times; the measurement was performed on Kimi-K2-Thinking at 32K/8K ISL/OSL. This addresses a key problem in agentic AI: agentic scenarios consume up to 15 times more tokens than traditional applications, and token cost is what limits their economics. In the second half of 2026, as Rubin arrives in Google Cloud, Lambda, and AWS, price competition for agentic inference will begin, and hyperscaler and neocloud procurement plans are already being built around this platform.
Why this matters for users
The new GPU generation is starting to actually work in clouds rather than remaining presentation material. Engineers and infrastructure operators should note that modular cable-free trays and 5-minute rack installation accelerate GPU generation changes in data centers. For those watching LLM inference prices, the benchmark is this: the claimed roughly 10x cheaper tokens on Rubin will start showing up in cloud prices with the first hyperscale deployments in the second half of 2026. Right now, the platform cannot be tried directly: there are no public APIs or prices for Rubin inference yet; access will appear first through Azure, then through other clouds, but agentic workflows that are not profitable at current prices should already be designed with the second half of 2026 in mind.
What is still unknown / limitations
All quantitative claims — a fourfold lower GPU requirement for training a 10-trillion-parameter MoE model and roughly 10x cheaper inference tokens — are vendor data: there are no independent benchmarks yet confirming the jump relative to GB200 NVL72. The inference measurement was performed by NVIDIA itself on a specific Kimi-K2-Thinking scenario (32K/8K ISL/OSL), and the methodology for the training claim has not been published — parallelism strategy, utilization, convergence trajectory. Public cloud prices for Rubin have not appeared yet, so the real price effect will only be visible with deployments in the second half of 2026, when the first independent measurements appear.
Sources
- Satya Nadella on X: photos of the first production Vera Rubin delivery to Microsoft data centers
- NVIDIA on X: statement on Vera Rubin entering full production
- NVIDIA Vera Rubin NVL72 — official product page
Author
Look at AI, editorial team
