GST-Bench tests whether VLMs develop global spatial awareness from video

ByteDance Seed

Research official 1 src. ~1 min

ByteDance Seed and academic collaborators introduce GST-Bench, a 2,762-question video VQA benchmark testing whether vision-language models can build global spatial awareness from long egocentric videos, including inference from novel viewpoints and mapping to top-down layouts. The best zero-shot VLM scores 42.68 versus a 79.08 human baseline; proprietary models fail mainly at cross-frame integration, open models fail at both local and global reasoning.

Why it matters

Quantifies a large, specific gap in VLM spatial cognition that current benchmarks mostly miss; 36 upvotes on Hugging Face Daily Papers.

Importance: 2/5

Notable benchmark quantifying a specific VLM weakness, below the 100-upvote bump threshold.

Sources