Nemotron 3 Ultra pipeline achieves IMO Gold with open-weight release
NVIDIA has released an open-source framework that enables its Nemotron 3 Ultra model to solve Olympiad-level mathematics problems at a gold-medal standard. The system utilizes a test-time compute pipeline that iteratively generates, verifies, and refines natural-language proofs without relying on formal provers or external tools.
Key points
- Nemotron 3 Ultra pipeline scored 30/42 at IMO 2026, meeting the gold-medal threshold.
- System uses iterative natural-language proof generation, verification, and refinement without formal tools.
- NVIDIA released two specialist checkpoints, training data, code, and a new 200-problem benchmark.
The approach involves training two specialist checkpoints from the base Nemotron 3 Ultra model using supervised fine-tuning and reinforcement learning. By combining these specialists with the general-availability model in an iterative search process, the system achieved a score of 30 out of 42 points at the International Mathematical Olympiad (IMO) 2026, surpassing the threshold for a gold medal.
To support further research, the team has released the post-trained checkpoints, training data, inference code, and submitted solutions. They also introduced Nemotron-IMO-Bench, a new benchmark consisting of 200 novel olympiad-level problems, providing the community with resources to evaluate and improve mathematical reasoning capabilities in large language models.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- LLM-Anchored Paralinguistic Boost for Alzheimer's Detection · 1 src
- New probability-wave framework links trader behavior to AGI architecture design · 1 src
- Cognitive Digital Twins: Self-Evolving Architectures · 1 src
- Linguistic Structure Enrichment Fails to Improve Text Coherence · 1 src
- New Methods Use Agent Internal States to Predict Success in Multi‑Turn Tasks · 1 src
Comments
via GitHub Discussions