Nvidia’s next-generation Vera Rubin platform has delivered its first major independent benchmark result, and the performance jump is large enough to create a different problem for AI companies: keeping up with Nvidia’s hardware cycle is becoming increasingly expensive.
In the latest MLPerf Inference v6.1 tests, Vera Rubin NVL72 delivered up to 3.7x the throughput of Nvidia’s GB300 NVL72 on the Qwen3-VL benchmark. On DeepSeek-R1, the improvement reached as much as 2.5x.
Nvidia detailed the results in its official MLPerf announcement, describing Vera Rubin’s entry as a preview submission ahead of broader deployment.
That distinction matters: 3.7x is the best measured improvement on a specific workload, not a claim that every AI application suddenly runs 3.7 times faster.
Nvidia Is Making Last Year’s AI Infrastructure Look Old Very Quickly
The benchmark also showed how rapidly the underlying economics are changing.
Nvidia reported that four GB300 NVL72 racks containing 288 GPUs achieved 99% scaling efficiency, while software improvements alone produced performance gains of as much as 1.6x over the previous MLPerf round.
AMD is pushing from the other direction. Its latest submission included Instinct MI355X and MI350-series accelerators, while a 512-GPU system operated by Crusoe produced the highest aggregate token throughput submitted in MLPerf history.
The result is an unusually fast replacement cycle. AI cloud operators are not simply adding GPUs; they are repeatedly deciding whether existing clusters remain competitive against newer systems capable of producing substantially more tokens per rack and per megawatt.
The Performance Race Is Becoming a Financing Race
That is where the story moves beyond Nvidia.
CoreWeave recently upsized a convertible-note offering from $3 billion to $3.7 billion, adding another large financing round to an industry already consuming enormous amounts of debt and equity. CoreWeave’s latest convertible financing showed how quickly GPU expansion is translating into balance-sheet pressure.
The same issue extends across the sector. Nebius has also turned to billions in financing, while the broader AI infrastructure boom is increasingly being funded with debt.
Bond investors are noticing. Reuters reported this week that AI-related borrowers are increasingly paying wider spreads as investors question how much additional capital the sector will require.
Nvidia argues that more throughput can lower the cost per token, meaning newer hardware may ultimately justify its investment through better economics. But Nvidia has not publicly provided a simple rack price that allows investors to compare the 3.7x benchmark gain directly with acquisition cost.