Alibaba cuts Qwen-Audio 3.1 prices by up to 95%, but TTS API is pending
Alibaba announced the Qwen-Audio 3.1 series at the Apsara Conference on September 23, 2026, introducing five audio models with significant price reductions. The company claims text-to-speech (TTS) costs are down about 70%, real-time conversation by 85%, and transcription (ASR) by up to 95%. However, a detailed check of official links reveals that only the ASR and Realtime models have active API…
Key points
- Alibaba released five Qwen-Audio 3.1 models, but only ASR and Realtime APIs are currently available.
- Realtime audio input costs $6.40 per million tokens, about one-fifth of OpenAI's equivalent model.
- TTS and TTS-Next models are not yet accessible via API, despite claims of 70% price cuts.
For developers, the Realtime model offers audio input at $6.40 per million tokens and output at $24, which is roughly one-fifth the price of OpenAI's gpt-realtime-2.1. The ASR pricing structure changed from per-second to per-token billing, making direct comparison with the previous version difficult without knowing the token conversion rate. While Alibaba claims its TTS model ranks first on Artificial Analysis, the new 3.1 version was not yet listed on the leaderboard at the time of verification, with the older 3.0 version holding third place.
The article notes that QwenCloud is operated by a Singaporean entity and advises caution regarding privacy, as voice data can identify individuals. Users are encouraged to test the free tier of the Realtime API to verify actual costs, as token-to-time conversion rates are not publicly documented.
Model pages: Qwen-Audio 3.1 → · Gemini 3.8 Flash TTS →
The story so far
2 episodes →- Alibaba cuts Qwen-Audio 3.1 prices by up to 95%, but TTS API is pendingthis story
Alibaba's audio AI 'Qwen-Audio 3.1' gets up to 95% price cut. After opening all the official price lists, only 2 out of 5 had API information
note.com · 26 September 2026
Loading the full article…
This text was published by note.com and written by オトモ。 とある音楽クリエイターの記録. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Anthropic releases Claude Opus 5.5 with 40% lower costs · 20 src
- Nvidia releases free Nemotron 3 Diarization model for real-time speaker identification · 1 src
- OpenAI says 80 to 90 percent of research targets GPT 7 and beyond · 1 src
- Meta opens early access for new Muse AI features · 3 src
- GPT-6 Astra and Claude Fable 5.1 compete in end‑to‑end software build · 2 src
Comments
via GitHub Discussions