DigestAI news desk

Cut through the AI noise.

Generative AI & Models4 min read

Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower prices

Alibaba’s Qwen team launched Qwen-Audio-3.1-Realtime, a full-duplex speech model designed for voice agents that can reason, call tools, and manage turn-taking in real-time conversations. The model is part of a five-model audio stack, including ASR (automatic speech recognition), TTS (text-to-speech), and real-time interaction capabilities. Pricing has dropped significantly:…

1 source

Key points

  • Alibaba’s Qwen-Audio-3.1-Realtime is a full-duplex speech model for voice agents, priced **85% lower** than before
  • Supports **262K context tokens**, function calling, web search, and multilingual speech recognition (14 languages)
  • Available only via **QwenCloud API**—no open weights; pricing starts at **$6.4 per 1M audio input tokens**

The model supports 262K context tokens (245K input, 16K output) and includes features like function calling, web search, structured outputs, and fine-tuning. Training was structured into three layers: Think (decision-making), Act (tool execution), and Speak (context-aware voice rendering). Benchmarks show improvements in multilingual accuracy (BBA average up from 81.7% to 88.1%) and reduced interference with background speech (dropping from 73% to 13% on Full-Duplex-Bench v1.5). However, GPT-Realtime-2 still leads in human red-team studies (96% vs. 92%). Pricing starts at $6.4 per 1M audio input tokens and $24 per 1M output tokens (text/audio).

Model page: Qwen-Audio-3.1-Realtime →

The story so far

3 episodes →
  1. Alibaba releases Qwen-Audio-3.1-Realtime for voice agents at 85% lower pricesthis story
Full story from MarkTechPost · by Asif RazzaqOpen source ↗

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

MarkTechPost · 29 September 2026

Loading the full article…

This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories