Tencent releases Ex-Omni, an 11B omni-modal model generating coordinated speech and 3D facial animation
Tencent
Tencent's HF org published weights of Ex-Omni, an 11B Qwen3-based omni-modal model that takes text or speech and outputs response text, speech units, and 52-dimensional facial blendshape coefficients for talking-face rendering. The accompanying paper (arXiv 2602.07106) was updated to v3 on Sep 3, 2026, introducing a blendshape-co-supervised speech-unit generator, a non-autoregressive blendshape decoder, and the 1.2M-sample InstructS2SF-1200K dataset.
Why it matters
It extends Chinese omni-modal LLM work beyond text and audio into joint 3D talking-face generation, an area the paper notes is largely unexplored; weights are on Tencent's official HF org (gated, EU access disallowed).
Importance: 2/5
Open-weights research artifact; niche but novel modality combination
Sources
official
tencent/Ex-Omni on Hugging Face