TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

Tencent

Research official + media 2 src. ~1 min

Omni-modal model for live-commerce that maps image, video, audio, and text into a unified representation space. Per-vGrid organizes timestamped tokens grouping each video grid with its temporally corresponding audio within explicit boundary tokens for temporal alignment. Three-stage supervised training progresses from omni-modal perception to instruction-following responses. A Faithful-RFT reinforcement stage scores final responses directly with task-verifiable feedback rather than reasoning-style rollout exploration. A scenario-oriented atomic-capability taxonomy and compact data production engine convert live-commerce streams into training signals for ASR, speaker analysis, product visual grounding, OCR, temporal grounding, video dense caption, and omni-modal QA. A synchronized length-grouped sampler reduces padding while preserving workloads across workers; a dynamic sampling strategy regenerates rollout groups with near-zero reward variance to maintain meaningful GRPO advantages.

Why it matters

HF Daily 55 upvotes on Aug 25. Tackles a noisy, multi-source real-world domain where product facts are scattered across speech, frames, product images, overlaid text, and queries — a useful blueprint for industry-specific omni-modal models.

Importance: 3/5

HF Daily 55 upvotes

Sources

official arXiv abstract