#vlm
- S²VOPD: self-supervised visual on-policy distillation lifts Qwen3.5-4B from 70.7% to 77.4% on six fine-grained perception benchmarks UC San Diego research
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Zhejiang University research
- Qwen releases Qwen-Drive-1.0, an open-weights vision-language model for autonomous driving Qwen/Alibaba models-llm
- DeepSeek open-sources V4-Flash-Vision-Exp, its first multimodal V4 model, under MIT DeepSeek models-llm
- OpenSearch-VL: Open Recipe for Training Frontier Multimodal Search Agents Tencent Hunyuan research
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs MCG-NJU (Nanjing University) research
- Yandex merges VLM and LLM into a single omnimodel powering Alice AI Yandex models-llm
- Astra: RL-Trained VLM Queries World Simulator for Spatial Reasoning research
- Wuhan AI Lab open-sources ZDTaichu5.0-9B, a spatial-reasoning VLM under 10B parameters Wuhan Artificial Intelligence Research Institute (Taichu) models-llm
- VRRL: Visually Grounded Self-Reflection for Vision-Language Models via RL UT Austin / Cornell research
- Yandex Smart Camera Gains Visual Q&A via Alice AI VLM Yandex tools
- Visual Contrastive Self-Distillation University of Maryland research
- GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? ByteDance Seed research
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL UC San Diego (Yunhao Yang, Nuno Vasconcelos, Yijiang Li et al.) research
- VKontakte Deploys LLM and VLM Models for In-Feed Product Recommendations VK AI tools