Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
Introduces a 1,253-example, expert-annotated benchmark across 60 subjects in four scientific disciplines for testing whether video generation models get scientific facts and causal dynamics right, not just visual realism. Finds models vary widely on scientific correctness with a pronounced gap between proprietary and open-source systems.
Why it matters
21 upvotes on HuggingFace Daily Papers; addresses a gap in video-gen evaluation that current benchmarks (visual quality only) miss.
Importance: 2/5
Notable paper (21 HF Daily Papers upvotes, below the 100-upvote bump threshold).
Sources
official
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
secondary
HuggingFace Daily Papers, 2026-08-11