VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
Tencent
A unified framework plus VWE-BENCH (2,616 assets, 323 seed worlds, 6,828 queries) and VibeWorlding-Gym (sandbox + rubric verifier) for training multimodal agents that infer intent, invoke 3D tools, and reflect on feedback. Even GPT-5.5 and Qwen3.8-Max score below 60%; RL-trained VibeWorlder-30B-A3B takes the best Pass@1 overall, beating closed-source models.
Why it matters
HF: 34 upvotes. First open benchmark plus training recipe showing open MLLMs can surpass frontier closed models on end-to-end 3D world construction.
Importance: 2/5
HF Daily 34 upvotes
Sources
official
arXiv:2608.15265