Alibaba Group has released Qwen-Audio-3.0-TTS, an innovative speech synthesis system that has already secured first place in the Artificial Analysis TTS Arena rankings. The model offers two operating modes: Flash for minimizing real-time latency and Plus for achieving the most natural sound.


What Happened
Alibaba has introduced the Qwen-Audio-3.0-TTS model, which supports 16 languages and 20 Chinese dialects. The system allows for control over emotional tone and speech tempo using 86 special tags (e.g., [laughing] or [whispers]) or via text instructions. The model is capable of generating long audio segments up to 3 minutes in a single pass and possesses high robustness to noise in source recordings during voice cloning.
Context
The development is production-oriented. A key technical solution was the implementation of a low-frequency tokenizer (12.5 Hz), which significantly reduces the computational cost of inference without critical loss of quality, providing a balance between performance and realism.
Why It Matters for the Industry
The release of this SOTA model sets a new standard for commercial voice AI assistants and voice cloning systems. The ability to work with long audio and the efficient architecture allow developers to quickly scale high-quality audio services while reducing infrastructure costs.
Why It Matters for Users
Users gain a tool for creating professional voiceovers with deep emotional expression. It is now possible to use simple text commands, such as "read like a bedtime story," to control speech style, as well as clone voices even from noisy audio recordings.
What Is Not Yet Known / Limitations
Current data does not specify the legal aspects and risks associated with the use of voice cloning technologies and privacy compliance.
Sources
- Qwen-Audio-3.0-TTS: Production-Oriented Speech Synthesis
- Alibaba Cloud Real-time speech synthesis guide
Author
Look at AI, Editorial Team
