OPRD: eliciting weak-to-strong generalization with on-policy reverse distillation
KAIST AI
On-Policy Reverse Distillation lets a stronger student surpass a weak teacher: it evaluates the teacher's policy shift on student rollouts and amplifies only the verifier-supported component of the student's policy gradient along that direction, preserving stationary points while accelerating learning beyond the teacher. With Aaron Courville among the authors.
Importance: 2/5
Alignment-relevant weak-to-strong result
Sources
official
arXiv 2609.08798