{"version":1,"type":"story","url":"https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower","json":"https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower.json","markdown":"https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower.md","slug":"alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower","headline":"Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices","summary":"Alibaba’s Qwen team launched **Qwen-Audio-3.1-Realtime**, a full-duplex speech model designed for voice agents that can reason, call tools, and manage turn-taking in real-time conversations. The model is part of a five-model audio stack, including ASR (automatic speech recognition), TTS (text-to-speech), and real-time interaction capabilities. Pricing has dropped significantly: **Qwen-Audio-3.1-Realtime** is **85% cheaper**, TTS is **70% cheaper**, and ASR is up to **95% cheaper** compared to previous versions. It is available exclusively via the **QwenCloud API** as *qwen-audio-3.1-realtime-plus* over WebSocket, with no open weights released.\n\nThe model supports **262K context tokens** (245K input, 16K output) and includes features like function calling, web search, structured outputs, and fine-tuning. Training was structured into three layers: **Think** (decision-making), **Act** (tool execution), and **Speak** (context-aware voice rendering). Benchmarks show improvements in multilingual accuracy (BBA average up from **81.7%** to **88.1%**) and reduced interference with background speech (dropping from **73%** to **13%** on Full-Duplex-Bench v1.5). However, GPT-Realtime-2 still leads in human red-team studies (**96%** vs. **92%**). Pricing starts at **$6.4 per 1M audio input tokens** and **$24 per 1M output tokens** (text/audio).","keyPoints":["Alibaba’s Qwen-Audio-3.1-Realtime is a full-duplex speech model for voice agents, priced **85% lower** than before","Supports **262K context tokens**, function calling, web search, and multilingual speech recognition (14 languages)","Available only via **QwenCloud API**—no open weights; pricing starts at **$6.4 per 1M audio input tokens**"],"whyItMatters":"This model advances real-time voice agent capabilities, lowering costs for developers building conversational AI tools. The **85% price cut** and **full-duplex turn-taking** could accelerate adoption in customer service, accessibility, and automation, though GPT’s lead in red-team resilience remains.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["Alibaba","Qwen","MarkTechPost","Marktechpost AI Media Inc."],"models":["Qwen-Audio-3.1-Realtime","Qwen-Audio-3.1-ASR-Flash-Filetrans","GPT-Realtime-2"],"people":["Asif Razzaq"]},"firstPublishedAt":"2026-09-29T04:58:15Z","updatedAt":"2026-09-29T04:58:15Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"MarkTechPost","title":"Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak","url":"https://marktechpost.com/2026/09/28/alibaba-qwen-releases-qwen-audio-3-1-realtime-a-full-duplex-voice-model-trained-to-think-act-and-decide-when-to-speak","publishedAt":"2026-09-29T04:58:15Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Alibaba Aggressively Cuts Qwen Audio Model Prices","url":"https://digestai.news/thread/nari-unveils-qwen3tts-and-qwen3asr-with-low-latency-and-competitive-pricing","storyCount":3},"cite":{"text":"Digest AI, \"Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices\", 29 September 2026, https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower","publisher":"Digest AI","title":"Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices","datePublished":"2026-09-29T04:58:15Z","url":"https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}