VGI-Bench: Probing Visual Intelligence in Video Generation Models
VGI-Bench packages 27 tasks / 810 instances under a two-level taxonomy of task domains and skill tags to probe zero-shot visual reasoning in generated video frames. The strongest evaluated generative system (Seedance 2.0) reaches only 51.0% under the proposed criteria.
Why it matters
HF Daily Papers Aug 27 at 143 upvotes; provides a principled yardstick for distinguishing world-model-like video generators from plausible-but-unreliable samplers, and quantifies how far current systems are from genuine visual reasoning.
Importance: 3/5
notable release; official+media confirmation; HF Daily 143 upvotes