DigestAI news desk

Cut through the AI noise.

Hardware & Compute4 min read

Open-source voice assistant Speakrail runs locally on an RTX 4090

Speakrail is a full-duplex voice assistant that runs entirely on a user’s machine, requiring only an RTX 4090 GPU. Released via GitHub, it supports real-time conversation, interruptions, and backchanneling like ‘mm-hmm’ without cutting off the user. The system combines Voxtral Mini 4B for speech recognition, Gemma 4 12B with a LoRA fine-tune for turn-taking logic, and Breeze TTS 2 for…

1 source primary source

Key points

  • Speakrail runs locally on an RTX 4090 GPU with ~75 GB disk space and Docker Compose
  • Achieves 94.0 on Full-Duplex-Bench v1.5, with median reply latency of 0.7–0.8 seconds
  • Default TTS (Breeze TTS 2) is non-commercial only; users must replace it for commercial use

The project requires ~75 GB of disk space and Docker Compose v2 with NVIDIA Container Toolkit. First-time setup takes 20 minutes to download models (~22 GB) and pre-render caches, while subsequent starts take a few minutes. Speakrail is licensed under Apache 2.0, with its base models (Gemma 4 12B, Voxtral Mini 4B) also under Apache 2.0. However, the default TTS (Breeze TTS 2) is restricted to non-commercial use only, requiring users to replace it for commercial applications. The project’s Full-Duplex-Bench scores reach 94.0, the highest among open-weight models, though written instructions are less reliable than spoken ones.

Full story from github.com · via Reddit AI communities primary sourceOpen source ↗

Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090

github.com · 5 October 2026

Loading the full article…

This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1source
Topics · follow one to build your own front page
Google DeepMindMistral AIBreezeBlueShugoAI LLCGemma 4 12BVoxtral Mini 4BBreeze TTS 2

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Hardware & Compute

All →

Related stories