OPRD: eliciting weak-to-strong generalization with on-policy reverse distillation

KAIST AI

Research official 2 src. ~1 min

On-Policy Reverse Distillation lets a stronger student surpass a weak teacher: it evaluates the teacher's policy shift on student rollouts and amplifies only the verifier-supported component of the student's policy gradient along that direction, preserving stationary points while accelerating learning beyond the teacher. With Aaron Courville among the authors.

Importance: 2/5

Alignment-relevant weak-to-strong result

Sources

official arXiv 2609.08798