Tencent open-sources AuK, a 1.5B speech foundation model for generation and editing
Tencent Hunyuan
Tencent released AuK and its distilled AuK-Flash variant under MIT license on Hugging Face: 1.5B-parameter speech foundation models that handle zero-shot TTS, voice cloning from a 10-second reference, instruct TTS from a text description, lyric rewriting in singing, and instruction-based editing of pitch, speed, emotion and timbre, plus denoising and source separation. Trained on ~3B instruction-audio instances and 1.95M hours across five task families; AuK-Flash runs fast 4-step inference, about 4.5x faster.
Why it matters
171 upvotes on HF Daily Papers. A single permissively licensed open model unifying TTS, cloning and natural-language speech editing is a strong free baseline for commercial voice applications, in both Chinese and English.
Importance: 4/5
Large open speech foundation model release + 171 upvotes on HF Daily Papers