# Open-source voice assistant Speakrail runs locally on an RTX 4090

Digest AI · Hardware & Compute · published 2026-10-05T15:50:04Z

Canonical: https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090

## Summary

Speakrail is a full-duplex voice assistant that runs entirely on a user’s machine, requiring only an RTX 4090 GPU. Released via GitHub, it supports real-time conversation, interruptions, and backchanneling like ‘mm-hmm’ without cutting off the user. The system combines **Voxtral Mini 4B** for speech recognition, **Gemma 4 12B** with a LoRA fine-tune for turn-taking logic, and **Breeze TTS 2** for text-to-speech, achieving a median latency of **0.7–0.8 seconds** per reply. It includes tools like weather checks, timers, and optional web search, with a focus on natural, uninterrupted dialogue flow.

The project requires **~75 GB** of disk space and **Docker Compose v2** with NVIDIA Container Toolkit. First-time setup takes **20 minutes** to download models (~22 GB) and pre-render caches, while subsequent starts take **a few minutes**. Speakrail is licensed under **Apache 2.0**, with its base models (**Gemma 4 12B**, **Voxtral Mini 4B**) also under Apache 2.0. However, the default TTS (**Breeze TTS 2**) is restricted to **non-commercial use only**, requiring users to replace it for commercial applications. The project’s **Full-Duplex-Bench** scores reach **94.0**, the highest among open-weight models, though written instructions are less reliable than spoken ones.

## Key points

- Speakrail runs locally on an RTX 4090 GPU with ~75 GB disk space and Docker Compose
- Achieves 94.0 on Full-Duplex-Bench v1.5, with median reply latency of 0.7–0.8 seconds
- Default TTS (Breeze TTS 2) is non-commercial only; users must replace it for commercial use

## Why it matters

A fully local, low-latency voice assistant could enable privacy-focused applications without cloud dependencies, appealing to developers and privacy-conscious users. Its real-time turn-taking and minimal latency may also advance conversational AI research for physical devices.

## Sources

1. [Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090](https://github.com/speakrail/speakrail) (github.com, 2026-10-05, primary source)

## Cite

Digest AI, "Open-source voice assistant Speakrail runs locally on an RTX 4090", 5 October 2026, https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/open-source-voice-assistant-speakrail-runs-locally-on-an-rtx-4090.json
