JD.com's Joy Future Academy releases JoyAI-Echo 1.5 — 5-minute audio-visual generation with cross-shot character and voice memory

JD.com (Joy Future Academy)

Video official 5 src. ~1 min

On 2026-08-28, JD's Echo Team at Joy Future Academy shipped JoyAI-Echo 1.5 (Echo-LongVideo) — a unified audio-visual generation system built on Lightricks LTX-2.3 with an 8-step DMD sampler, paired cross-modal memory (image + audio slots) for character appearance and voice identity, and Gemma 3 (12B IT) as the text encoder. The reference-to-video pipeline takes a text prompt, an optional first-frame condition, and up to seven ordered reference memory slots per shot, and ships with consumer-GPU profiles via layer-wise DiT offload and tiled Video VAE decoding. Weights are released in BF16, FP8, and FP4 precisions under the LTX-2 Community License on the main branch of the open-source repo.

Why it matters

First open-weights long-horizon audio-visual model at minute-scale with explicit character/voice memory slots — directly targets the 'persistent story' use case that Veo 3.1, Sora-class, and Wan 3.0 leave unsolved, and pairs with a Director Agent for orchestrated multi-shot pipelines.

Importance: 5/5

flagship release; official confirmation; 5 sources

Sources