News · High impactBack to News

Alibaba Launches Qwen-Audio-3.0-TTS: Tops Global Speech Quality Leaderboard with Production-Grade Semantic Controllability and Ultra-Low Latency

Alibaba's Tongyi Lab has launched Qwen-Audio-3.0-TTS, a hosted text-to-speech system that topped the Artificial Analysis Speech Arena leaderboard with a 1,238 Elo score. Offered via Alibaba Cloud Model Studio as an API-only service, it features a Plus tier for high-quality dubbing and a Flash tier with 300ms latency for voice agents. It supports fine-grained control via natural language and 86 inline tags, across 16 languages and 20 Chinese dialect regions.

Tongyi Lab Releases Qwen-Audio-3.0-TTS and Tops Speech Quality Leaderboard

Alibaba's Tongyi Lab released the production-oriented, hosted text-to-speech (TTS) system Qwen-Audio-3.0-TTS in July 2026. The model claimed the #1 position on the independent Artificial Analysis Speech Arena leaderboard with a Quality Elo rating of 1,238, surpassing Google's Gemini 3.1 Flash TTS, SpeechifyAI's Simba 3.2, and Cartesia's Sonic 3.5. This milestone demonstrates that Chinese voice generation models have reached global tier-one standards in overall naturalness, timbre fidelity, and end-to-end output quality.

Hosted Dual-Tier Architecture: Catering to High-Fidelity Dubbing and Real-Time Voice Agents

Distinguished from Alibaba's previously released open-source Qwen3-TTS series (0.6B/1.7B models for local deployment), the newly launched Qwen-Audio-3.0-TTS is a proprietary, API-only hosted service delivered via the Alibaba Cloud Model Studio (Bailian). To target different industrial use cases, Alibaba offers two distinct service tiers:

  • Plus Version: Optimized for voice naturalness and timbre fidelity, making it ideal for audiobooks, video narrations, and professional dubbing;
  • Flash Version: Prioritizes response speeds, reducing first-packet latency to approximately 300 milliseconds to empower highly interactive real-time voice agents and virtual assistants.

Core Technical Innovations: Low-Frame-Rate Tokenizer and 86 Fine-Grained Inline Control Tags

Architecturally, the model utilizes a 12.5 Hz low-frame-rate speech tokenizer alongside a five-stage progressive training paradigm, displaying high robustness against noisy references and supporting high-fidelity, single-pass synthesis for up to 3 minutes. Qwen-Audio-3.0-TTS introduces a "production-level" control paradigm: developers can customize voice styles using natural language commands or apply 86 inline tags for localized tuning. The model currently supports 16 languages and 20 Chinese dialect regions.

Next step

Keep tracking Alibaba Cloud

Continue along the same topic.

Open entity record