{"version":1,"type":"story","url":"https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090","json":"https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090.json","markdown":"https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090.md","slug":"open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090","headline":"Open-source voice assistant Speakrail runs locally on an RTX 4090","summary":"Speakrail is a full-duplex voice assistant that runs entirely on a user’s machine, requiring only an RTX 4090 GPU. Released via GitHub, it supports real-time conversation, interruptions, and backchanneling like ‘mm-hmm’ without cutting off the user. The system combines **Voxtral Mini 4B** for speech recognition, **Gemma 4 12B** with a LoRA fine-tune for turn-taking logic, and **Breeze TTS 2** for text-to-speech, achieving a median latency of **0.7–0.8 seconds** per reply. It includes tools like weather checks, timers, and optional web search, with a focus on natural, uninterrupted dialogue flow.\n\nThe project requires **~75 GB** of disk space and **Docker Compose v2** with NVIDIA Container Toolkit. First-time setup takes **20 minutes** to download models (~22 GB) and pre-render caches, while subsequent starts take **a few minutes**. Speakrail is licensed under **Apache 2.0**, with its base models (**Gemma 4 12B**, **Voxtral Mini 4B**) also under Apache 2.0. However, the default TTS (**Breeze TTS 2**) is restricted to **non-commercial use only**, requiring users to replace it for commercial applications. The project’s **Full-Duplex-Bench** scores reach **94.0**, the highest among open-weight models, though written instructions are less reliable than spoken ones.","keyPoints":["Speakrail runs locally on an RTX 4090 GPU with ~75 GB disk space and Docker Compose","Achieves 94.0 on Full-Duplex-Bench v1.5, with median reply latency of 0.7–0.8 seconds","Default TTS (Breeze TTS 2) is non-commercial only; users must replace it for commercial use"],"whyItMatters":"A fully local, low-latency voice assistant could enable privacy-focused applications without cloud dependencies, appealing to developers and privacy-conscious users. Its real-time turn-taking and minimal latency may also advance conversational AI research for physical devices.","category":{"slug":"hardware","name":"Hardware & Compute","url":"https://digestai.news/category/hardware"},"entities":{"companies":["Google DeepMind","Mistral AI","BreezeBlue","ShugoAI LLC"],"models":["Gemma 4 12B","Voxtral Mini 4B","Breeze TTS 2"],"people":[]},"firstPublishedAt":"2026-10-05T15:50:04Z","updatedAt":"2026-10-05T15:50:04Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090","url":"https://github.com/speakrail/speakrail","publishedAt":"2026-10-05T15:50:04Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wyc1sc/speakrail_a_lowlatency_fullylocal_voice_assistant/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"Open-source voice assistant Speakrail runs locally on an RTX 4090\", 5 October 2026, https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090","publisher":"Digest AI","title":"Open-source voice assistant Speakrail runs locally on an RTX 4090","datePublished":"2026-10-05T15:50:04Z","url":"https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}