TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Reformulates vision-language-action modeling as a direct V+L to action mapping instead of routing through a large language-model backbone, reaching 97.7% average success on LIBERO with only 0.2B parameters, 31.2ms latency, and under 1GB VRAM on a consumer RTX 4090.
Why it matters
Reached 122 upvotes on HuggingFace Daily Papers for 2026-07-30; matches or beats much larger LLM-centric VLA systems while making real-time robot control feasible on commodity hardware.
Importance: 4/5
Notable research paper with >=100 HF Daily Papers upvotes (122), triggering the confidence heuristic bump.
Sources
official
TurboVLA