Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Microsoft
A 4B-parameter codec-native multimodal model whose custom visual tokenizer (Mage-ViT) selectively encodes motion- and residual-rich regions of video streams instead of uniform frame sampling, cutting visual token count by over 75% and giving up to 3.5x faster inference.
Why it matters
At 4B parameters it matches Qwen3-VL-4B on general VQA and surpasses the much larger Phi-4-R-V-15B on spatial and video understanding, addressing real-time perception as a known weak spot for VLMs.
Importance: 3/5
Notable efficiency-focused multimodal research paper from Microsoft, below the 100-upvote heuristic threshold.
Sources
official
Mage-VL