DAPD identifies 'privilege illusion' failure mode in on-policy self-distillation
Shanghai AI Laboratory
Identifies a 'privilege illusion' failure mode in on-policy self-distillation, where a student model learns behavior dependent on training-time privileged information it cannot access at inference. Proposes Dual-Path and Dual-Source Anchoring to align reference and rollout behavior bidirectionally, improving over prior on-policy self-distillation by roughly 2-2.8 points across Qwen3 models from 4B to 32B.
Why it matters
Reported with over 100 upvotes on HuggingFace Daily Papers; addresses a subtle but broadly applicable failure mode in the increasingly common on-policy self-distillation training recipe.
Importance: 3/5
HF Daily Papers with >=100 upvotes bump applied.