Nari unveils Qwen3‑TTS and Qwen3‑ASR with low latency and competitive pricing
Nari announced two new speech models, Qwen3‑TTS for text‑to‑speech and Qwen3‑ASR for automatic speech recognition. Both run on optimized API endpoints that promise enterprise‑grade latency and cost efficiency.
Key points
- Qwen3‑TTS and Qwen3‑ASR launched with API endpoints optimized for low latency and cost.
- Benchmarks show the TTS beats ElevenLabs and the ASR beats Deepgram in speed and price.
- Nari offers free beta, early‑access pricing, and enterprise deployment options including dedicated infrastructure.
Independent measurements cited in the post show the TTS endpoint beating ElevenLabs on latency and price per million characters, while the ASR endpoint outperforms Deepgram on speed and hourly cost. The services are in a free public beta with early‑access pricing, and Nari says the stack uses custom kernels, batching, quantization and routing to hit the “Pareto frontier” on benchmark suites.
Nari also promotes dedicated or private deployment options and fine‑tuning capabilities, positioning the models for large‑scale production use. The announcement references the team behind Dia and Narvatar, suggesting a broader ecosystem for multimodal AI workloads.
Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
narilabs.com · 8 September 2026
Model APIs
Optimized model APIs
Optimized multimodal model endpoints built to be fastest in production: starting with speech.
Learn More
Realtime TTS
LatencyElevenLabs vs Nari FastTTFA
CostElevenLabs vs Nari StandardUSD / 1M characters
Streaming STT
LatencyDeepgram vs Nari FastTTFS
CostDeepgram vs Nari StandardUSD / hour of audio
Based on preliminary independent measurements of publicly available TTS and STT endpoints from US East. Results may vary by region, connection, and workload. Free endpoints are best effort and carry no latency or uptime SLO. Free Public Beta is available for a limited time; Early Access pricing is shown. Pricing checked Sep 8, 2026 from Artificial Analysis.
Inference
Dedicated Multimodal Inference for Enterprise
Model and workload-specific runtimes deliver predictable latency and production-scale serving across managed, dedicated, and private infrastructure.
Learn More Model-specific optimizations
Kernels, batching, quantization, and routing tuned to your workload.
Proven on our stack
The same technology we used to achieve the Pareto frontier and #1 on Benchmarks.
Deployment flexibility
Managed, dedicated, or private infrastructure.
Training
Finetune a model for your use-case
Fine-tune multimodal models on data that matches your use-case and deploy using our optimized stack. Work with the team behind Dia and Narvatar.
This text was published by narilabs.com and written by Nari Labs. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1 source- Hacker News discussion · 44 points news.ycombinator.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Generative AI & Models
All →- Anthropic's Models Hack External Companies · 2 src
- OpenAI's GPT-6 Astra tops ErdosBench math benchmark despite no math focus · 4 src
- OpenAI Releases GPT‑6 Astra to ChatGPT Plus Subscribers · 32 src
- Meta Launches Personal AI Agent for Mass Adoption · 2 src
- Google Tests Payment for Content in AI Search Features · 1 src
Comments
via GitHub Discussions