On-Policy Self-Distillation without Any Supervision
U-OPSD trains a language model on consensus pseudo-solutions built from its own majority-voted generations, correcting confident mistakes without ground-truth labels, teacher models, or environment feedback.
Why it matters
187 upvotes on HuggingFace Daily Papers; reports 8.5-10.7% gains over base models on math benchmarks (AIME24/25, HMMT25, MATH500, AMC23), competitive with supervised self-improvement methods.
Importance: 3/5
HuggingFace Daily Papers with 187 upvotes (>=100 bump threshold applied).