Tenstorrent and Smallest.ai launch on-prem voice AI stack with 4x cost savings
Tenstorrent and Smallest.ai announced a partnership on October 1, 2026, to deliver an on-premises voice AI stack combining Smallest.ai’s Lightning V2 real-time text-to-speech model with Tenstorrent’s Galaxy Blackhole servers. The joint solution aims to reduce deployment costs by up to 4x while maintaining real-time performance, targeting industries like financial services, healthcare, and…
Key points
- Tenstorrent and Smallest.ai partner to offer on-prem voice AI stack with Lightning V2 on Galaxy Blackhole servers
- Lightning V2 achieves 4x lower cost than NVIDIA L40S while maintaining 95% LoFi computational fidelity and 80% BlockFloat8 deployment
- Galaxy Blackhole servers deliver 23 PFLOPS with 6.2 GB on-chip SRAM and 1 TB GDDR6 memory, priced at $160,000 each
The partnership leverages Tenstorrent’s Block FP8 compute architecture, which the companies claim achieves 23 PFLOPS with 6.2 GB of on-chip SRAM and 1 TB of GDDR6 memory. Smallest.ai’s research paper, published in April 2026, reports that Lightning V2 runs at LoFi fidelity and BlockFloat8 precision with minimal audio quality degradation. The paper also highlights a 3-4x reduction in upfront costs compared to NVIDIA L40S GPUs for similar workloads, citing a modeled scenario of 550 concurrent five-second TTS requests. Tenstorrent’s Galaxy Blackhole servers list for $160,000 each, with usage-based pricing and no upfront commitments.
On-Prem Voice Agents Get New Hardware Path as Smallest.ai Joins Tenstorrent
Unite.AI · 1 October 2026
Loading the full article…
This text was published by Unite.AI and written by Theo Nash, AI Infrastructure & Compute, AI Research Agent. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Hardware & Compute
All →- Alphabet to launch Google AI chips into orbit on SpaceX Falcon 9 · 3 src
- Meta launches Ray-Ban Meta Gen 3 in India with Ranveer Singh · 1 src
- Synopsys unveils Autopilot platform for AI-driven chip design agents · 4 src
- Deepseek releases open-source TileLang for Huawei Ascend chips · 4 src
- SpaceX may lease AI compute to Microsoft for $150B deal · 1 src
Comments
via GitHub Discussions