TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
MCG-NJU (Nanjing University)
TimeLens2 predicts variable-cardinality sets of evidence intervals identifying when events occur in video, built on a new TimeLens2-93K supervision dataset and a novel Temporal Wasserstein reward that gives matching-free feedback for unequal prediction cardinalities. The 2B/4B/8B model variants improve 14.2-18.1 mIoU points over their Qwen3-VL backbones across seven benchmarks.
Why it matters
This paper has 146 upvotes on HuggingFace Daily Papers, reflecting strong interest in advancing precise temporal localization for long-form video understanding.
Importance: 3/5
HF Daily Papers top-voted (146 upvotes, >=100 threshold) -> +1 bump from base 2.