Editable Visual Design: coding agents generate layered HTML/CSS visuals instead of flattened images

Research official 2 src. ~1 min

A team around Junyan Ye proposes a new visual-generation paradigm where a coding agent uses a VLM as the creative brain and an image model as an on-demand visual-world simulator: the agent generates isolated assets, writes native HTML/CSS, and refines the design against rendered feedback. The result is editable, layer-decoupled artifacts, in contrast to end-to-end diffusion output (GPT-Image-2, Nano Banana) which yields flattened bitmaps with error-prone text.

Why it matters

Top-upvoted paper of the Sep 5 HF chart (42 upvotes); points at a post-image-editor direction where designs stay structured and editable rather than baked into pixels.

Importance: 2/5

Top of a weekend HF Daily Papers chart, below the 100-upvote bar

Sources

secondary Editable Visual Design — Hugging Face Daily Papers (syndicated from arXiv)