🤖 Tencent Hunyuan open-sources 1.5B parameter speech model AuK
On September 9, 2026, the company released the source code, MIT-licensed weights, and technical report (arXiv:2609.08936). The model understands natural language commands: TTS and voice cloning from a reference, TTS from a description, word replacement and deletion in recordings, re-singing text in a song while preserving the melody, emotion and timbre changes, denoising, speaker separation, and vocal extraction.
🌍 Open weights, code, ComfyUI nodes, Python API, and a fine-tuning pipeline with day-0 support in SGLang-Omni provide an open-source alternative to closed TTS tools.
👤 Demos are already available on HuggingFace and ModelScope: the tencent/AuK and tencent/AuK-Flash weights can be downloaded to clone voices locally, rewrite words, and clean up noise. Russian is not officially supported, but the multilingual Qwen2.5-Omni-3B encoder leaves room for hope.
Source 1: https://github.com/Tencent-Hunyuan/AuK Source 2: https://huggingface.co/tencent/AuK
