TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

MCG-NJU (Nanjing University)

Research official 1 src. ~1 min

TimeLens2 predicts variable-cardinality sets of evidence intervals identifying when events occur in video, built on a new TimeLens2-93K supervision dataset and a novel Temporal Wasserstein reward that gives matching-free feedback for unequal prediction cardinalities. The 2B/4B/8B model variants improve 14.2-18.1 mIoU points over their Qwen3-VL backbones across seven benchmarks.

Why it matters

This paper has 146 upvotes on HuggingFace Daily Papers, reflecting strong interest in advancing precise temporal localization for long-form video understanding.

Importance: 3/5

HF Daily Papers top-voted (146 upvotes, >=100 threshold) -> +1 bump from base 2.

Sources