# Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices

Digest AI · Generative AI & Models · published 2026-09-29T04:58:15Z

Canonical: https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower

## Summary

Alibaba’s Qwen team launched **Qwen-Audio-3.1-Realtime**, a full-duplex speech model designed for voice agents that can reason, call tools, and manage turn-taking in real-time conversations. The model is part of a five-model audio stack, including ASR (automatic speech recognition), TTS (text-to-speech), and real-time interaction capabilities. Pricing has dropped significantly: **Qwen-Audio-3.1-Realtime** is **85% cheaper**, TTS is **70% cheaper**, and ASR is up to **95% cheaper** compared to previous versions. It is available exclusively via the **QwenCloud API** as *qwen-audio-3.1-realtime-plus* over WebSocket, with no open weights released.

The model supports **262K context tokens** (245K input, 16K output) and includes features like function calling, web search, structured outputs, and fine-tuning. Training was structured into three layers: **Think** (decision-making), **Act** (tool execution), and **Speak** (context-aware voice rendering). Benchmarks show improvements in multilingual accuracy (BBA average up from **81.7%** to **88.1%**) and reduced interference with background speech (dropping from **73%** to **13%** on Full-Duplex-Bench v1.5). However, GPT-Realtime-2 still leads in human red-team studies (**96%** vs. **92%**). Pricing starts at **$6.4 per 1M audio input tokens** and **$24 per 1M output tokens** (text/audio).

## Key points

- Alibaba’s Qwen-Audio-3.1-Realtime is a full-duplex speech model for voice agents, priced **85% lower** than before
- Supports **262K context tokens**, function calling, web search, and multilingual speech recognition (14 languages)
- Available only via **QwenCloud API**—no open weights; pricing starts at **$6.4 per 1M audio input tokens**

## Why it matters

This model advances real-time voice agent capabilities, lowering costs for developers building conversational AI tools. The **85% price cut** and **full-duplex turn-taking** could accelerate adoption in customer service, accessibility, and automation, though GPT’s lead in red-team resilience remains.

## Sources

1. [Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak](https://marktechpost.com/2026/09/28/alibaba-qwen-releases-qwen-audio-3-1-realtime-a-full-duplex-voice-model-trained-to-think-act-and-decide-when-to-speak) (MarkTechPost, 2026-09-29)

Part of the developing story: [Alibaba Aggressively Cuts Qwen Audio Model Prices](https://digestai.news/thread/nari-unveils-qwen3tts-and-qwen3asr-with-low-latency-and-competitive-pricing) (3 stories)

## Cite

Digest AI, "Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices", 29 September 2026, https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/alibaba-releases-qwen-audio-3-1-realtime-for-voice-agents-at-85-lower.json
