GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

ByteDance Seed

Research official 1 src. ~1 min

Introduces a video question-answering benchmark built from nearly 6,800 minutes of synthetic video to test whether vision-language models can integrate spatial observations across a video into a coherent global representation, testing 22 state-of-the-art models.

Why it matters

Fifth most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 42 upvotes; finds the best zero-shot model scores only 42.68 versus a human score of 79.08, a large capability gap in spatial video reasoning.

Importance: 2/5

Notable HuggingFace Daily Paper (research) — Fifth most-upvoted paper on HuggingFace Daily Papers for 2026-08-09 with 42 upvotes; finds the best zero-shot model scores only 42.68 versus a human score of 79.08, a large capability gap in spatial video reasoning.

Sources