LLaDA-Image: building strong image generators with fully open training recipes
inclusionAI (Ant Group)
A 6B Diffusion Transformer image generator trained from scratch, paired with a frozen vision-language module built on the LLaDA2.0-Mini diffusion language model, trained on a 220M-sample pipeline with RMSNorm and the Muon optimizer. The visual generative prior is built through image-only pre-training and mid-training rather than paired image-text data; a distilled Turbo variant generates in 2-4 steps. Claims SOTA among open models on Qwen-Image-Bench (53.53 EN / 53.38 ZH) with weights and full training recipes released.
Why it matters
226 upvotes on HF Daily Papers Sep 4 — the first fully open, from-scratch training recipe for a competitive open image generator
Importance: 4/5
notable paper (3) + ≥100 upvotes on HF Daily Papers (+1)
Sources
official
HF Daily Papers — LLaDA-Image