JD.com's Joy Future Academy releases JoyAI-Echo 1.5 — 5-minute audio-visual generation with cross-shot character and voice memory
JD.com (Joy Future Academy)
On 2026-08-28, JD's Echo Team at Joy Future Academy shipped JoyAI-Echo 1.5 (Echo-LongVideo) — a unified audio-visual generation system built on Lightricks LTX-2.3 with an 8-step DMD sampler, paired cross-modal memory (image + audio slots) for character appearance and voice identity, and Gemma 3 (12B IT) as the text encoder. The reference-to-video pipeline takes a text prompt, an optional first-frame condition, and up to seven ordered reference memory slots per shot, and ships with consumer-GPU profiles via layer-wise DiT offload and tiled Video VAE decoding. Weights are released in BF16, FP8, and FP4 precisions under the LTX-2 Community License on the main branch of the open-source repo.
Why it matters
First open-weights long-horizon audio-visual model at minute-scale with explicit character/voice memory slots — directly targets the 'persistent story' use case that Veo 3.1, Sora-class, and Wan 3.0 leave unsolved, and pairs with a Director Agent for orchestrated multi-shot pipelines.
Importance: 5/5
flagship release; official confirmation; 5 sources