DigestAI news desk
Generative AI & Models updated 1 min read

Nari unveils Qwen3‑TTS and Qwen3‑ASR with low latency and competitive pricing

Nari announced two new speech models, Qwen3‑TTS for text‑to‑speech and Qwen3‑ASR for automatic speech recognition. Both run on optimized API endpoints that promise enterprise‑grade latency and cost efficiency.

1 source HN 44

Key points

  • Qwen3‑TTS and Qwen3‑ASR launched with API endpoints optimized for low latency and cost.
  • Benchmarks show the TTS beats ElevenLabs and the ASR beats Deepgram in speed and price.
  • Nari offers free beta, early‑access pricing, and enterprise deployment options including dedicated infrastructure.

Independent measurements cited in the post show the TTS endpoint beating ElevenLabs on latency and price per million characters, while the ASR endpoint outperforms Deepgram on speed and hourly cost. The services are in a free public beta with early‑access pricing, and Nari says the stack uses custom kernels, batching, quantization and routing to hit the “Pareto frontier” on benchmark suites.

Nari also promotes dedicated or private deployment options and fine‑tuning capabilities, positioning the models for large‑scale production use. The announcement references the team behind Dia and Narvatar, suggesting a broader ecosystem for multimodal AI workloads.

Full story from narilabs.com · by Nari Labs · via Hacker News Open source ↗

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

narilabs.com · 8 September 2026

Model APIs

Optimized model APIs

Optimized multimodal model endpoints built to be fastest in production: starting with speech.

Learn More

Realtime TTS

LatencyElevenLabs vs Nari FastTTFA

CostElevenLabs vs Nari StandardUSD / 1M characters

Streaming STT

LatencyDeepgram vs Nari FastTTFS

CostDeepgram vs Nari StandardUSD / hour of audio

Based on preliminary independent measurements of publicly available TTS and STT endpoints from US East. Results may vary by region, connection, and workload. Free endpoints are best effort and carry no latency or uptime SLO. Free Public Beta is available for a limited time; Early Access pricing is shown. Pricing checked Sep 8, 2026 from Artificial Analysis.

Inference

Dedicated Multimodal Inference for Enterprise

Model and workload-specific runtimes deliver predictable latency and production-scale serving across managed, dedicated, and private infrastructure.

Learn More Model-specific optimizations

Kernels, batching, quantization, and routing tuned to your workload.

Proven on our stack

The same technology we used to achieve the Pareto frontier and #1 on Benchmarks.

Deployment flexibility

Managed, dedicated, or private infrastructure.

Training

Finetune a model for your use-case

Fine-tune multimodal models on data that matches your use-case and deploy using our optimized stack. Work with the team behind Dia and Narvatar.

This text was published by narilabs.com and written by Nari Labs. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1 source
Topics · follow one to build your own front page
NariElevenLabsDeepgramDiaNarvatarQwen3-TTSQwen3-ASR

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories