Dream-RSI: recursive self-improvement through evolving worlds
Turns an agent's accumulated discovery history into a replay simulator and trains the exploration policy in it ('dreaming'), giving cheap off-policy feedback instead of costly online rollouts. The refined policy is redeployed to make new discoveries that expand the simulator, closing a recursive self-improvement loop as a non-invasive orchestration layer over an unchanged coding agent.
Why it matters
263 upvotes on HF Daily Papers; shows competitive or better discovery quality at lower cost across algorithm engineering, math optimization and GPU kernel engineering.
Importance: 4/5
Notable paper + 263 upvotes on HF Daily Papers (+1 bump)