{"version":1,"type":"story","url":"https://digestai.news/story/kyutai-releases-voice-of-reason-speech-to-speech-model-for-spoken-math","json":"https://digestai.news/story/kyutai-releases-voice-of-reason-speech-to-speech-model-for-spoken-math.json","markdown":"https://digestai.news/story/kyutai-releases-voice-of-reason-speech-to-speech-model-for-spoken-math.md","slug":"kyutai-releases-voice-of-reason-speech-to-speech-model-for-spoken-math","headline":"Kyutai releases Voice of Reason speech-to-speech model for spoken math","summary":"Kyutai announced two open‑weight speech‑to‑speech models, called Voice of Reason, built on the GLM‑4‑Voice‑9B foundation. The models solve math problems directly from spoken input, removing the need for a transcription step or a separate text LLM. On the spoken GSM8K benchmark the direct model reaches 70.3% accuracy, and with a reasoning‑chunk variant (Stitch) it climbs to 77.1%, up from the base model’s 27.3%.\n\nTraining used 150,616 math problems from Orca‑Math, which were rewritten by Qwen3‑235B and voiced by Kyutai’s DSM TTS system. Supervised fine‑tuning lifted accuracy to 61.7%, and a second reinforcement‑learning stage—sampling four replies per question and scoring them with a Qwen3‑235B‑A22B‑2507 judge—added the remaining gains. The RL run involved 1,500 updates on 16 H100 GPUs. Both 9B checkpoints run on a single H100 for inference, are released under the GLM‑4‑Voice license, and are currently not hosted by Hugging Face’s inference service.\n\nThe work demonstrates that speech‑native models can reason about math without the latency of cascaded pipelines, preserving paralinguistic cues and keeping interaction fluid. Kyutai’s results suggest reinforcement learning can substantially improve reasoning performance in audio‑only systems.","keyPoints":["Voice of Reason raises spoken GSM8K accuracy from 27.3% to 77.1% using reinforcement learning.","Training used 150,616 math problems, 16 H100 GPUs, and 1,500 RL updates.","Both 9B checkpoints are open‑weight, run on a single H100, but not hosted on Hugging Face inference."],"whyItMatters":"Speech‑native reasoning cuts latency and avoids transcription errors, opening interactive AI applications that can solve problems aloud in real time.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["Kyutai","Hugging Face"],"models":["GLM-4-Voice-9B","Voice of Reason-9b","Voice of Reason-stitch-9b","Qwen3-235B","Qwen3-235B-A22B-2507","GSM8K"],"people":[]},"firstPublishedAt":"2026-09-23T06:33:18Z","updatedAt":"2026-09-23T06:33:18Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"MarkTechPost","title":"Kyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning","url":"https://marktechpost.com/2026/09/22/kyutai-releases-voice-of-reason-a-speech-native-model-that-solves-spoken-math-with-reinforcement-learning","publishedAt":"2026-09-23T06:33:18Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Kyutai releases Voice of Reason speech-to-speech model for spoken math\", 23 September 2026, https://digestai.news/story/kyutai-releases-voice-of-reason-speech-to-speech-model-for-spoken-math","publisher":"Digest AI","title":"Kyutai releases Voice of Reason speech-to-speech model for spoken math","datePublished":"2026-09-23T06:33:18Z","url":"https://digestai.news/story/kyutai-releases-voice-of-reason-speech-to-speech-model-for-spoken-math"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}