Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices
Alibaba’s Qwen team launched Qwen-Audio-3.1-Realtime, a full-duplex speech model designed for voice agents that can reason, call tools, and manage turn-taking in real-time conversations. The model is part of a five-model audio stack, including ASR (automatic speech recognition), TTS (text-to-speech), and real-time interaction capabilities. Pricing has dropped significantly:…
Key points
- Alibaba’s Qwen-Audio-3.1-Realtime is a full-duplex speech model for voice agents, priced **85% lower** than before
- Supports **262K context tokens**, function calling, web search, and multilingual speech recognition (14 languages)
- Available only via **QwenCloud API**—no open weights; pricing starts at **$6.4 per 1M audio input tokens**
The model supports 262K context tokens (245K input, 16K output) and includes features like function calling, web search, structured outputs, and fine-tuning. Training was structured into three layers: Think (decision-making), Act (tool execution), and Speak (context-aware voice rendering). Benchmarks show improvements in multilingual accuracy (BBA average up from 81.7% to 88.1%) and reduced interference with background speech (dropping from 73% to 13% on Full-Duplex-Bench v1.5). However, GPT-Realtime-2 still leads in human red-team studies (96% vs. 92%). Pricing starts at $6.4 per 1M audio input tokens and $24 per 1M output tokens (text/audio).
Model page: Qwen-Audio-3.1-Realtime →
The story so far
3 episodes →- Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower pricesthis story
Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak
MarkTechPost · 29 September 2026
Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- OpenAI scraps GPT-6.1 Astra release over safety concerns · 9 src
- Anthropic launches Claude Opus 5.5 with 40% lower costs and Fable 5.1-level performance · 9 src
- Anthropic to launch Sonnet 5.5 next week, targeting cost savings · 9 src
- xAI releases Grok 4.7 on Amazon Bedrock · 1 src
- Study finds AI chatbots offer narrower knowledge than Google search · 1 src
Comments
via GitHub Discussions