Dream-RSI: recursive self-improvement through evolving worlds

Google

Research official 2 src. ~1 min

Turns an agent's accumulated discovery history into a replay simulator and trains the exploration policy in it ('dreaming'), giving cheap off-policy feedback instead of costly online rollouts. The refined policy is redeployed to make new discoveries that expand the simulator, closing a recursive self-improvement loop as a non-invasive orchestration layer over an unchanged coding agent.

Why it matters

263 upvotes on HF Daily Papers; shows competitive or better discovery quality at lower cost across algorithm engineering, math optimization and GPU kernel engineering.

Importance: 4/5

Notable paper + 263 upvotes on HF Daily Papers (+1 bump)

Sources

official Dream-RSI — Hugging Face Daily Papers (syndicated from arxiv.org)