Tongyi Lab Releases Qwen-Audio-3.0-TTS and Tops Speech Quality Leaderboard
Alibaba's Tongyi Lab released the production-oriented, hosted text-to-speech (TTS) system Qwen-Audio-3.0-TTS in July 2026. The model claimed the #1 position on the independent Artificial Analysis Speech Arena leaderboard with a Quality Elo rating of 1,238, surpassing Google's Gemini 3.1 Flash TTS, SpeechifyAI's Simba 3.2, and Cartesia's Sonic 3.5. This milestone demonstrates that Chinese voice generation models have reached global tier-one standards in overall naturalness, timbre fidelity, and end-to-end output quality.
Hosted Dual-Tier Architecture: Catering to High-Fidelity Dubbing and Real-Time Voice Agents
Distinguished from Alibaba's previously released open-source Qwen3-TTS series (0.6B/1.7B models for local deployment), the newly launched Qwen-Audio-3.0-TTS is a proprietary, API-only hosted service delivered via the Alibaba Cloud Model Studio (Bailian). To target different industrial use cases, Alibaba offers two distinct service tiers:
- Plus Version: Optimized for voice naturalness and timbre fidelity, making it ideal for audiobooks, video narrations, and professional dubbing;
- Flash Version: Prioritizes response speeds, reducing first-packet latency to approximately 300 milliseconds to empower highly interactive real-time voice agents and virtual assistants.
Core Technical Innovations: Low-Frame-Rate Tokenizer and 86 Fine-Grained Inline Control Tags
Architecturally, the model utilizes a 12.5 Hz low-frame-rate speech tokenizer alongside a five-stage progressive training paradigm, displaying high robustness against noisy references and supporting high-fidelity, single-pass synthesis for up to 3 minutes. Qwen-Audio-3.0-TTS introduces a "production-level" control paradigm: developers can customize voice styles using natural language commands or apply 86 inline tags for localized tuning. The model currently supports 16 languages and 20 Chinese dialect regions.