Wuhan AI Lab open-sources ZDTaichu5.0-9B, a spatial-reasoning VLM under 10B parameters

Wuhan Artificial Intelligence Research Institute (Taichu)

Models / LLM official + media 3 src. ~1 min

Taichu released ZDTaichu5.0-9B, an open-weight multimodal model pairing a Qwen3.5-9B language backbone with a C-RADIOv4-H vision encoder for images, video, and 128K-context text. It claims leading spatial-reasoning scores among sub-10B VLMs (SparBench, ViewSpatial, MMSI-Bench) plus agent results (TAU2-Bench 87.7) and uses entropy-gated adaptive recurrent reasoning to allocate extra latent computation to hard tokens; weights and an FP8 build are on Hugging Face with code on GitHub.

Why it matters

Bundles spatial reasoning, embodied-AI grounding, and tool-use into a single open small model — a practical base for robotics and on-device agents rather than a vision-only specialist.

Importance: 2/5

Notable small open-weights VLM release with strong spatial-reasoning claims

Sources