Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

Microsoft

Research official + media 2 src. ~1 min

A 4B-parameter codec-native multimodal model whose custom visual tokenizer (Mage-ViT) selectively encodes motion- and residual-rich regions of video streams instead of uniform frame sampling, cutting visual token count by over 75% and giving up to 3.5x faster inference.

Why it matters

At 4B parameters it matches Qwen3-VL-4B on general VQA and surpasses the much larger Phi-4-R-V-15B on spatial and video understanding, addressing real-time perception as a known weak spot for VLMs.

Importance: 3/5

Notable efficiency-focused multimodal research paper from Microsoft, below the 100-upvote heuristic threshold.

Sources

official Mage-VL