Rethinking On-Policy Distillation of LLMs II: near-full gains from a single training example

Research official 2 src. ~1 min

Pushes on-policy distillation to its data-minimal limit: training on a single query keeps improving for hundreds of steps and recovers most of full-data OPD gains. The mechanism is state coverage — one query already reaches 71.5% of the states full-data OPD visits, and 16 diverse queries reach 98.9%. Conclusion: OPD is 'data-overfed but algorithm-starved' — step efficiency, not data quantity, is the real bottleneck.

Why it matters

reframes distillation data budgets: near-full gains from a handful of queries

Importance: 2/5

notable follow-up paper; below the 100-upvote bump threshold

Sources