WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Tencent ARC Lab

Research official 2 src. ~1 min

A video world model with a camera-queryable implicit 3D-aware memory: the target viewpoint conditions how historical observations are compressed into the generator's token budget, without explicit depth correspondences. A jointly-trained memory encoder plus pose-conditioned readout and few-step distillation support minute-scale streaming scene exploration from a single image or text prompt.

Why it matters

Top-voted paper of the day on HF Daily Papers (194 upvotes) and a step toward interactive open-world exploration with long-horizon consistency, the core blocker for playable world models.

Importance: 4/5

Top HF Daily Papers of the day (194 upvotes, >=100 bump)

Sources