Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis




Vera Rubin NVL72 is the second generation of Nvidia’s rack-scale Oberon architecture, and its gains on inference come from extreme co-design. Early results from engineering samples are encouraging. Vera Rubin NVL72 running DeepSeek R1 delivers 5.4x performance per MW and 5x performance per dollar over GB200 NVL72 today, and the gap is even wider against GB200 NVL72 during its early bringup in 2025. Vera Rubin is still in the early bringup stage now, so we expect the gap to continue widen. Rubin's inference performance will keep improving as software matures, the same pattern we demonstrated for Blackwell in our InferenceX benchmarks, and Rubin still has a long runway ahead.Nvidia has also recently made available their first public release of the Rubin (SM_107) software stack with CUDA13…