DigestAI news desk

Vera Rubin NVL72 Delivers 67× More Tokens per Dollar in Agentic Inference

Vera Rubin NVL72, NVIDIA’s new co‑designed accelerator, has shown a 67‑fold increase in tokens per dollar on the AgentX agentic inference benchmark, outperforming the GB300 and Blackwell platforms. Early pre‑release software already achieves up to 7× higher token throughput per megawatt and more than double the profit per gigawatt, according to SemiAnalysis’s analysis.

1 source

Key points

  • Rubin NVL72 achieves 67× more tokens per dollar than GB300 on AgentX benchmark.
  • Performance per watt reaches up to 7× higher token throughput compared to Blackwell.
  • At 80 TPS, Rubin delivers 16× more tokens per 3‑year rental TCO than Blackwell.

The benchmark, widely validated by major cloud buyers such as Google Cloud, Microsoft Azure, Oracle, Meta, and supported by the ML community (vLLM, LMCache, SGLang, PyTorch, Huggingface), demonstrates that Rubin can generate 18–39× more tokens per dollar at realistic interactivity targets (80–120 TPS). With a 3‑year rental cost of $8.5/hr/chip versus $5/hr/chip for Blackwell, Rubin also delivers up to 16× more tokens per rental TCO and 39% higher annual revenue per gigawatt.

Key figures include 61% higher P90 interactivity, 62% more tokens per rental TCO at 80 TPS, and a projected $159.5 B annual revenue per gigawatt at 75 TPS interactivity. The platform’s DSX MaxLPS dynamic power shifting further boosts datacenter density.

The story so far

7 episodes →
  1. Vera Rubin NVL72 Delivers 67× More Tokens per Dollar in Agentic Inference this story
Read the original at SemiAnalysis · by Bryan Shan Open source ↗
Topics · follow one to build your own front page
NVIDIAGoogle CloudMicrosoft AzureOracleMetaOpenAIRubin NVL72GB300BlackwellDeepSeek V4 ProDeepSeek V4.1 FlashMI355XJensen HuangIan BuckNick ComlyKedar PotdarRohit Nagraj

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Hardware & Compute

All →

Related stories