Tencent releases Ex-Omni, an 11B omni-modal model generating coordinated speech and 3D facial animation

Tencent

Research official 2 src. ~1 min

Tencent's HF org published weights of Ex-Omni, an 11B Qwen3-based omni-modal model that takes text or speech and outputs response text, speech units, and 52-dimensional facial blendshape coefficients for talking-face rendering. The accompanying paper (arXiv 2602.07106) was updated to v3 on Sep 3, 2026, introducing a blendshape-co-supervised speech-unit generator, a non-autoregressive blendshape decoder, and the 1.2M-sample InstructS2SF-1200K dataset.

Why it matters

It extends Chinese omni-modal LLM work beyond text and audio into joint 3D talking-face generation, an area the paper notes is largely unexplored; weights are on Tencent's official HF org (gated, EU access disallowed).

Importance: 2/5

Open-weights research artifact; niche but novel modality combination

Sources