Survey: scaling reasoning models beyond human supervision via a five-level autonomy ladder

Research official 2 src. ~1 min

A 72-page position paper by 19 authors frames how large reasoning models can keep improving as human oversight fades, organizing progress along a reward axis (human judgments to reusable automated verifiers) and an experience axis (human-designed tasks to self-generated curricula). It proposes an L0-L4 ladder of learning-loop autonomy and names the growing risks: reward hacking, feedback drift, curriculum collapse, and environment errors.

Why it matters

15 upvotes on HF Daily Papers Sep 1; gives the field a shared vocabulary for the shift from human-labeled data to self-sustaining training loops.

Importance: 2/5

Notable position paper on HF Daily Papers (15 upvotes), official sources

Sources