Alibaba launches Qwen Audio 3.1 with five speech models and price cuts
Alibaba’s Qwen AI team released Qwen Audio 3.1, a suite of five models covering speech‑to‑text (ASR), an enhanced ASR‑Next, text‑to‑speech (TTS), a diffusion‑based TTS‑Next, and a real‑time interaction model. The ASR model improves multilingual and dialect recognition while automatically removing filler words and repetitions. ASR‑Next adds multi‑speaker identification with timestamps and can…
Key points
- Qwen Audio 3.1 adds five models for ASR, ASR‑Next, TTS, TTS‑Next, and real‑time interaction.
- Price cuts: ASR up to 95 %, TTS about 70 %, real‑time roughly 85 % lower.
- New features include multi‑speaker ID with timestamps, emotion detection, and diffusion‑based voice‑plus‑sound generation.
The TTS model supports multilingual synthesis and cross‑language voice transfer, letting users steer emotion, speed, and style through simple prompts such as “Read this with a sharp, commanding tone, demanding respect.” TTS‑Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass. The real‑time model enables simultaneous speaking and listening, instant interruption, and, according to Qwen, slows its response and adds empathy when it senses a low mood. Alibaba also announced price reductions: ASR up to 95 % off, TTS about 70 % off, and the real‑time model roughly 85 % off. Further details are on the Qwen blog and Qwen Cloud.
Model page: Qwen Audio 3.1 →
The story so far
3 episodes →- Alibaba launches Qwen Audio 3.1 with five speech models and price cutsthis story
Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
The Decoder · 23 September 2026
Loading the full article…
This text was published by The Decoder and written by Matthias Bastian. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Anthropic releases Claude Opus 5.5 with $4 input and $20 output pricing, claiming 40% lower cost than Opus 5 · 22 src
- Anthropic releases Claude Opus 5.5 and OpenAI counters with cheaper GPT-6 Sol and Luna models · 28 src
- Opinion: OpenAI could copy TypeSafe's Jev classifier and embed it in future models · 3 src
- Kyutai releases Voice of Reason speech-to-speech model for spoken math · 1 src
- Anthropic releases Claude Opus 5.5 with Fable 5.1-level performance and 40% lower cost · 12 src
Comments
via GitHub Discussions