TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

Research official + media 2 src. ~1 min

Reformulates vision-language-action modeling as a direct V+L to action mapping instead of routing through a large language-model backbone, reaching 97.7% average success on LIBERO with only 0.2B parameters, 31.2ms latency, and under 1GB VRAM on a consumer RTX 4090.

Why it matters

Reached 122 upvotes on HuggingFace Daily Papers for 2026-07-30; matches or beats much larger LLM-centric VLA systems while making real-time robot control feasible on commodity hardware.

Importance: 4/5

Notable research paper with >=100 HF Daily Papers upvotes (122), triggering the confidence heuristic bump.

Sources

official TurboVLA