Editable Visual Design: coding agents generate layered HTML/CSS visuals instead of flattened images
A team around Junyan Ye proposes a new visual-generation paradigm where a coding agent uses a VLM as the creative brain and an image model as an on-demand visual-world simulator: the agent generates isolated assets, writes native HTML/CSS, and refines the design against rendered feedback. The result is editable, layer-decoupled artifacts, in contrast to end-to-end diffusion output (GPT-Image-2, Nano Banana) which yields flattened bitmaps with error-prone text.
Why it matters
Top-upvoted paper of the Sep 5 HF chart (42 upvotes); points at a post-image-editor direction where designs stay structured and editable rather than baked into pixels.
Importance: 2/5
Top of a weekend HF Daily Papers chart, below the 100-upvote bar
Sources
official
Editable Visual Design (arXiv)