🤖 Alibaba has released Qwen-Audio-3.0-TTS — a new model for speech synthesis

Alibaba Group has introduced the Qwen-Audio-3.0-TTS system, which has secured the top spot in the Artificial Analysis TTS Arena rankings. The model is available in a 'Flash' version for real-time applications and a 'Plus' version for maximum naturalness. The system supports 16 languages and allows for emotional control via 86 specialized tags.

🌍 The release of this SOTA model sets a new standard for commercial AI assistants. The use of a low-frequency tokenizer (12.5 Hz) significantly reduces inference costs.

👤 Users can now create high-quality voiceovers with emotional expression and clone voices from noisy recordings.

Source 1: https://funaudiollm.github.io/qwen-audio-3.0-tts/ Source 2: https://www.alibabacloud.com/help/en/model-studio/realtime-tts-user-guide