Show-Harness: just a VLM agent can play robots

Show Lab, NUS

Research official 1 src. ~1 min

Show-Harness is an Embodied Harness that lets vision-language models control robots by reasoning over discrete semantic action units that deterministic interpreters compile into low-level robot actions. The same interface enables zero-shot control by closed-source frontier VLMs, cheap adaptation of small open VLMs, and GUI-based demonstration collection without teleoperation hardware, outperforming representative agentic and VLA baselines.

Why it matters

126 upvotes on HuggingFace Daily Papers (Sep 10). Evidence for the 'harness over model' thesis in robotics: the right interface unlocks embodied capability from general VLMs without embodiment-specific pretraining.

Importance: 4/5

Notable paper with 126 upvotes on HF Daily (bump for ≥100 upvotes)

Sources