GST-Bench tests whether VLMs develop global spatial awareness from video
ByteDance Seed
ByteDance Seed and academic collaborators introduce GST-Bench, a 2,762-question video VQA benchmark testing whether vision-language models can build global spatial awareness from long egocentric videos, including inference from novel viewpoints and mapping to top-down layouts. The best zero-shot VLM scores 42.68 versus a 79.08 human baseline; proprietary models fail mainly at cross-frame integration, open models fail at both local and global reasoning.
Why it matters
Quantifies a large, specific gap in VLM spatial cognition that current benchmarks mostly miss; 36 upvotes on Hugging Face Daily Papers.
Importance: 2/5
Notable benchmark quantifying a specific VLM weakness, below the 100-upvote bump threshold.