{"version":1,"type":"story","url":"https://digestai.news/story/alibaba-launches-qwen-audio-3-1-with-five-speech-models-and-price-cuts","json":"https://digestai.news/story/alibaba-launches-qwen-audio-3-1-with-five-speech-models-and-price-cuts.json","markdown":"https://digestai.news/story/alibaba-launches-qwen-audio-3-1-with-five-speech-models-and-price-cuts.md","slug":"alibaba-launches-qwen-audio-3-1-with-five-speech-models-and-price-cuts","headline":"Alibaba launches Qwen Audio 3.1 with five speech models and price cuts","summary":"Alibaba’s Qwen AI team released Qwen Audio 3.1, a suite of five models covering speech‑to‑text (ASR), an enhanced ASR‑Next, text‑to‑speech (TTS), a diffusion‑based TTS‑Next, and a real‑time interaction model. The ASR model improves multilingual and dialect recognition while automatically removing filler words and repetitions. ASR‑Next adds multi‑speaker identification with timestamps and can detect emotions, ambient sounds, and machine noise.\n\nThe TTS model supports multilingual synthesis and cross‑language voice transfer, letting users steer emotion, speed, and style through simple prompts such as “Read this with a sharp, commanding tone, demanding respect.” TTS‑Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass. The real‑time model enables simultaneous speaking and listening, instant interruption, and, according to Qwen, slows its response and adds empathy when it senses a low mood. Alibaba also announced price reductions: ASR up to 95 % off, TTS about 70 % off, and the real‑time model roughly 85 % off. Further details are on the Qwen blog and Qwen Cloud.","keyPoints":["Qwen Audio 3.1 adds five models for ASR, ASR‑Next, TTS, TTS‑Next, and real‑time interaction.","Price cuts: ASR up to 95 %, TTS about 70 %, real‑time roughly 85 % lower.","New features include multi‑speaker ID with timestamps, emotion detection, and diffusion‑based voice‑plus‑sound generation."],"whyItMatters":"Much cheaper, multilingual speech models let developers embed high‑quality voice interfaces in apps, games, and customer‑service tools, expanding AI audio adoption.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["Alibaba"],"models":["Qwen Audio 3.1","ASR-Next","TTS-Next"],"people":[]},"firstPublishedAt":"2026-09-23T12:31:20Z","updatedAt":"2026-09-23T12:31:20Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"The Decoder","title":"Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent","url":"https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent","publishedAt":"2026-09-23T12:31:20Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Alibaba's Multimodal AI Model Race","url":"https://digestai.news/thread/alibaba-releases-qwen3-8-omni-flash-a-1m-token-omni-modal-model-with-agentic","storyCount":3},"cite":{"text":"Digest AI, \"Alibaba launches Qwen Audio 3.1 with five speech models and price cuts\", 23 September 2026, https://digestai.news/story/alibaba-launches-qwen-audio-3-1-with-five-speech-models-and-price-cuts","publisher":"Digest AI","title":"Alibaba launches Qwen Audio 3.1 with five speech models and price cuts","datePublished":"2026-09-23T12:31:20Z","url":"https://digestai.news/story/alibaba-launches-qwen-audio-3-1-with-five-speech-models-and-price-cuts"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}