Alibaba's Multimodal AI Model Race
The saga follows Alibaba's rollout of increasingly advanced multimodal AI models, from the Qwen-Image-2.1 system that claims superiority over closed‑source image generators to the new Qwen3.8‑Omni‑Flash, a 1‑million‑token omni‑modal model with agentic audio‑video capabilities. With the latest Omni‑Flash launch, Alibaba now positions itself at the forefront of open‑source, large‑scale, agentic AI, challenging industry leaders across vision, language, and audio‑video domains.
-
Alibaba's Qwen-Image-2.1 claims to beat closed models in image generation
Alibaba's Qwen AI team released Qwen-Image-2.1, an open‑weight image generation and editing model with 7 billion parameters. The team says it outperforms most closed models on its own benchmark,…
1 source -
Alibaba releases Qwen3.8-Omni-Flash, a 1M-token omni-modal model with agentic audio‑video
Alibaba’s Qwen team announced the launch of Qwen3.8-Omni-Flash, its first omni‑modal model that handles text, images, audio and video and returns text. The model runs as a hosted API on QwenCloud,…
3 sources