{"version":1,"type":"story","url":"https://digestai.news/story/tenstorrent-and-smallest-ai-launch-on-prem-voice-ai-stack-with-4x-cost","json":"https://digestai.news/story/tenstorrent-and-smallest-ai-launch-on-prem-voice-ai-stack-with-4x-cost.json","markdown":"https://digestai.news/story/tenstorrent-and-smallest-ai-launch-on-prem-voice-ai-stack-with-4x-cost.md","slug":"tenstorrent-and-smallest-ai-launch-on-prem-voice-ai-stack-with-4x-cost","headline":"Tenstorrent and Smallest.ai launch on-prem voice AI stack with 4x cost savings","summary":"Tenstorrent and Smallest.ai announced a partnership on October 1, 2026, to deliver an on-premises voice AI stack combining Smallest.ai’s Lightning V2 real-time text-to-speech model with Tenstorrent’s Galaxy Blackhole servers. The joint solution aims to reduce deployment costs by up to 4x while maintaining real-time performance, targeting industries like financial services, healthcare, and telecom that prioritize data sovereignty and high-volume workloads.\n\nThe partnership leverages Tenstorrent’s Block FP8 compute architecture, which the companies claim achieves 23 PFLOPS with 6.2 GB of on-chip SRAM and 1 TB of GDDR6 memory. Smallest.ai’s research paper, published in April 2026, reports that Lightning V2 runs at LoFi fidelity and BlockFloat8 precision with minimal audio quality degradation. The paper also highlights a 3-4x reduction in upfront costs compared to NVIDIA L40S GPUs for similar workloads, citing a modeled scenario of 550 concurrent five-second TTS requests. Tenstorrent’s Galaxy Blackhole servers list for $160,000 each, with usage-based pricing and no upfront commitments.","keyPoints":["Tenstorrent and Smallest.ai partner to offer on-prem voice AI stack with Lightning V2 on Galaxy Blackhole servers","Lightning V2 achieves 4x lower cost than NVIDIA L40S while maintaining 95% LoFi computational fidelity and 80% BlockFloat8 deployment","Galaxy Blackhole servers deliver 23 PFLOPS with 6.2 GB on-chip SRAM and 1 TB GDDR6 memory, priced at $160,000 each"],"whyItMatters":"This partnership enables organizations to deploy high-performance, real-time voice agents on-premises at a fraction of the cost of traditional GPU solutions, addressing data sovereignty concerns in sensitive industries like finance and healthcare.","category":{"slug":"hardware","name":"Hardware & Compute","url":"https://digestai.news/category/hardware"},"entities":{"companies":["Tenstorrent","Smallest.ai"],"models":["Lightning V2","Lightning V3"],"people":["Amr Elashmawi","Ranjith M S","Akshat Mandloi","Sudarshan Kamath"]},"firstPublishedAt":"2026-10-01T17:09:01Z","updatedAt":"2026-10-01T17:09:01Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"Unite.AI","title":"On-Prem Voice Agents Get New Hardware Path as Smallest.ai Joins Tenstorrent","url":"https://unite.ai/on-prem-voice-agents-get-new-hardware-path-as-smallest-ai-joins-tenstorrent","publishedAt":"2026-10-01T17:09:01Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Tenstorrent and Smallest.ai launch on-prem voice AI stack with 4x cost savings\", 1 October 2026, https://digestai.news/story/tenstorrent-and-smallest-ai-launch-on-prem-voice-ai-stack-with-4x-cost","publisher":"Digest AI","title":"Tenstorrent and Smallest.ai launch on-prem voice AI stack with 4x cost savings","datePublished":"2026-10-01T17:09:01Z","url":"https://digestai.news/story/tenstorrent-and-smallest-ai-launch-on-prem-voice-ai-stack-with-4x-cost"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}