Nvidia releases Nemotron 3 Diarization, a 100M-parameter model tracking up to 8 speakers
Nvidia announced the open‑weight speaker diarization model Nemotron 3 Diarization on Hugging Face. The 100 million‑parameter model can identify who spoke when in audio streams, handling up to eight overlapping speakers with a single checkpoint that works for both offline recordings and real‑time streaming. The weights are released under the OpenMDW License 1.1, which permits commercial use, and…
Key points
- Nemotron 3 Diarization is a 100M‑parameter open‑weight model that tracks up to eight overlapping speakers.
- The model ranked first on Voice Arena’s Diarization‑Bench with 14.72% DER, about 24% better than the next system.
- Weights are released under the OpenMDW License 1.1, allowing commercial use, and run on Nvidia GPUs via NeMo.
In Voice Arena’s initial Diarization‑Bench, Nemotron 3 Diarization ranked first among twelve systems, achieving a 14.72% diarization error rate (DER) versus 19.3% for the runner‑up, roughly a 24% relative improvement. Nvidia notes these results may change once Voice Arena completes its Version 1 evaluation. The model is not yet available through Hugging Face Inference Providers, but production deployments can use Baseten, DigitalOcean, or Argmax Pro SDK 3.
Model page: Nemotron 3 Diarization →
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
MarkTechPost · 23 September 2026
Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Anthropic releases Claude Opus 5.5 and OpenAI counters with cheaper GPT-6 Sol and Luna models · 32 src
- Anthropic releases Claude Opus 5.5 with $4 input and $20 output pricing, claiming 40% lower cost than Opus 5 · 27 src
- Google launches Gemini 3.8 Flash and Flash‑Lite text‑to‑speech models · 5 src
- Anthropic gives Claude users one free usage reset token · 1 src
- Alibaba's Qwen-Image-2.1 claims to beat closed models in image generation · 3 src
Comments
via GitHub Discussions