GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
ByteDance Seed
Introduces a video question-answering benchmark built from nearly 6,800 minutes of synthetic video to test whether vision-language models can integrate spatial observations across a video into a coherent global representation, testing 22 state-of-the-art models.
Why it matters
Fifth most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 42 upvotes; finds the best zero-shot model scores only 42.68 versus a human score of 79.08, a large capability gap in spatial video reasoning.
Importance: 2/5
Notable HuggingFace Daily Paper (research) — Fifth most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 42 upvotes; finds the best zero-shot model scores only 42.68 versus a human score of 79.08, a large capability gap in spatial video reasoning.