Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Introduces JoyAI-Echo-1.5, a unified audio-visual generation system with two variants — long-form narrative video and an interactive 6-DoF world model — built on composable cross-shot memory, speech-filtered audio cues, and a causal few-step generator trained with progressive teacher forcing and Self-Gradient Forcing.
Why it matters
HF Daily Papers Aug 27 at 1.96k upvotes — by far the highest-voted new paper in the window; demonstrates that long-horizon audio-visual consistency with controllable camera trajectories is now within reach for unified generation systems.
Importance: 4/5
flagship release; official+media confirmation; HF Daily 1.96k upvotes