SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

ByteDance

Research official 1 src. ~1 min

ByteDance's model for multi-speaker expressive speech and audio generation supports both instruct (caption-driven) and zero-shot (reference-audio-driven) synthesis, combining reward-conditioned quality control, Engram conditioning, and a unified MoE trained with curriculum learning plus GRPO post-training. Targets applications like dubbing, audio drama, and short-video production.

Why it matters

HuggingFace Daily Papers listing with 142 upvotes; part of ByteDance's SwanAIGC audio research line alongside SwanVoice and SwanBench-Speech.

Importance: 4/5

Notable paper with 142 HF Daily Papers upvotes (>=100 threshold), bumped +1.

Sources