Open-source voice assistant Speakrail runs locally on an RTX 4090
Speakrail is a full-duplex voice assistant that runs entirely on a user’s machine, requiring only an RTX 4090 GPU. Released via GitHub, it supports real-time conversation, interruptions, and backchanneling like ‘mm-hmm’ without cutting off the user. The system combines Voxtral Mini 4B for speech recognition, Gemma 4 12B with a LoRA fine-tune for turn-taking logic, and Breeze TTS 2 for…
Key points
- Speakrail runs locally on an RTX 4090 GPU with ~75 GB disk space and Docker Compose
- Achieves 94.0 on Full-Duplex-Bench v1.5, with median reply latency of 0.7–0.8 seconds
- Default TTS (Breeze TTS 2) is non-commercial only; users must replace it for commercial use
The project requires ~75 GB of disk space and Docker Compose v2 with NVIDIA Container Toolkit. First-time setup takes 20 minutes to download models (~22 GB) and pre-render caches, while subsequent starts take a few minutes. Speakrail is licensed under Apache 2.0, with its base models (Gemma 4 12B, Voxtral Mini 4B) also under Apache 2.0. However, the default TTS (Breeze TTS 2) is restricted to non-commercial use only, requiring users to replace it for commercial applications. The project’s Full-Duplex-Bench scores reach 94.0, the highest among open-weight models, though written instructions are less reliable than spoken ones.
Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090
github.com · 5 October 2026Loading the full article…
This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Hardware & Compute
All →- AMD benchmarks Gorgon Halo AI chips against Intel, ahead of RTX Spark launch · 1 src
- ResearchAndMarkets adds GPUaaS forecast: 27.2% CAGR to $118.8B by 2036 · 1 src
- Nvidia partners with Microsoft and Valar to speed up US nuclear builds · 1 src
- Huawei's Atlas 960 Super Pod cuts power use by over 550 kW per machine · 1 src
- Nvidia raises Shield TV Pro price by $100 to $299.99 due to AI-driven component costs · 5 src
Comments
via GitHub Discussions