SenseNova-U1.5: native unified visual intelligence without encoders or VAEs
SenseTime
SenseTime's SenseNova-U1.5 is an 8B mixture-of-transformers model doing visual understanding, reasoning, generation and editing in one encoder-free, VAE-free architecture — no separate vision encoder, no diffusion VAE. It uses spatially coherent patch reconstruction as the visual interface, native resolution up to 4K, and post-training with specialized experts (aesthetics, bilingual text rendering, infographics, editing) merged via multi-expert on-policy distillation. The report shows gains in fidelity, text rendering, multi-reference editing and identity-preserving edits; training code for SFT, RL and on-policy distillation is planned for open release.
Why it matters
139 upvotes on HuggingFace Daily Papers (Sep 11); a credible open-weight push toward fully native unified multimodality, removing the encoder/VAE scaffolding most unified models still carry.
Importance: 4/5
Notable paper with 139 upvotes on HF Daily (bump for >=100 upvotes)