Survey: scaling reasoning models beyond human supervision via a five-level autonomy ladder
A 72-page position paper by 19 authors frames how large reasoning models can keep improving as human oversight fades, organizing progress along a reward axis (human judgments to reusable automated verifiers) and an experience axis (human-designed tasks to self-generated curricula). It proposes an L0-L4 ladder of learning-loop autonomy and names the growing risks: reward hacking, feedback drift, curriculum collapse, and environment errors.
Why it matters
15 upvotes on HF Daily Papers Sep 1; gives the field a shared vocabulary for the shift from human-labeled data to self-sustaining training loops.
Importance: 2/5
Notable position paper on HF Daily Papers (15 upvotes), official sources