SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
ByteDance
ByteDance's model for multi-speaker expressive speech and audio generation supports both instruct (caption-driven) and zero-shot (reference-audio-driven) synthesis, combining reward-conditioned quality control, Engram conditioning, and a unified MoE trained with curriculum learning plus GRPO post-training. Targets applications like dubbing, audio drama, and short-video production.
Why it matters
HuggingFace Daily Papers listing with 142 upvotes; part of ByteDance's SwanAIGC audio research line alongside SwanVoice and SwanBench-Speech.
Importance: 4/5
Notable paper with 142 HF Daily Papers upvotes (>=100 threshold), bumped +1.