LLaDA-Image: building strong image generators with fully open training recipes

inclusionAI (Ant Group)

Research official 2 src. ~1 min

A 6B Diffusion Transformer image generator trained from scratch, paired with a frozen vision-language module built on the LLaDA2.0-Mini diffusion language model, trained on a 220M-sample pipeline with RMSNorm and the Muon optimizer. The visual generative prior is built through image-only pre-training and mid-training rather than paired image-text data; a distilled Turbo variant generates in 2-4 steps. Claims SOTA among open models on Qwen-Image-Bench (53.53 EN / 53.38 ZH) with weights and full training recipes released.

Why it matters

226 upvotes on HF Daily Papers Sep 4 — the first fully open, from-scratch training recipe for a competitive open image generator

Importance: 4/5

notable paper (3) + ≥100 upvotes on HF Daily Papers (+1)

Sources