TPUv7 Ironwood beats NVIDIA B200/B300 on inference cost by up to 50%
SemiAnalysis has released the first third-party benchmark results for Google’s TPUv7 Ironwood, marking a significant shift as the accelerator moves beyond internal use to compete directly with NVIDIA in external inference markets. In apples-to-apples comparisons using FP8 precision, Ironwood delivers up to 50% better performance per dollar than NVIDIA’s B200 and B300 GPUs. At a standard…
Key points
- TPUv7 Ironwood offers up to 50% better performance per dollar than NVIDIA B200/B300 in FP8 inference benchmarks.
- New TorchTPU backend enables native PyTorch support, replacing the older TorchAX translation layer for improved stability.
- Anthropic committed to over one million TPUs, driving external demand and validating TPU economics outside Google Cloud.
The benchmarks utilize the new TorchTPU backend, a native PyTorch integration that replaces the previous TorchAX translation layer, aiming to improve stability and performance for open-weight models like Qwen3.5 397B. While NVIDIA GPUs currently lead in FP4 precision due to TPUv7’s lack of native support, Google anticipates TPUv8i will close this gap. Anthropic is a key driver of this externalization, having committed to over one million TPUs for training and inference. The software stack is expected to leave private beta and be open-sourced around October, with future optimizations targeting disaggregated serving and agentic workloads.
The story so far
5 episodes →- TPUv7 Ironwood beats NVIDIA B200/B300 on inference cost by up to 50% this story
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
Comments
via GitHub Discussions