Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

Research official + media 2 src. ~1 min

Introduces JoyAI-Echo-1.5, a unified audio-visual generation system with two variants — long-form narrative video and an interactive 6-DoF world model — built on composable cross-shot memory, speech-filtered audio cues, and a causal few-step generator trained with progressive teacher forcing and Self-Gradient Forcing.

Why it matters

HF Daily Papers Aug 27 at 1.96k upvotes — by far the highest-voted new paper in the window; demonstrates that long-horizon audio-visual consistency with controllable camera trajectories is now within reach for unified generation systems.

Importance: 4/5

flagship release; official+media confirmation; HF Daily 1.96k upvotes

Sources