DAPD identifies 'privilege illusion' failure mode in on-policy self-distillation

Shanghai AI Laboratory

Research official 1 src. ~1 min

Identifies a 'privilege illusion' failure mode in on-policy self-distillation, where a student model learns behavior dependent on training-time privileged information it cannot access at inference. Proposes Dual-Path and Dual-Source Anchoring to align reference and rollout behavior bidirectionally, improving over prior on-policy self-distillation by roughly 2-2.8 points across Qwen3 models from 4B to 32B.

Why it matters

Reported with over 100 upvotes on HuggingFace Daily Papers; addresses a subtle but broadly applicable failure mode in the increasingly common on-policy self-distillation training recipe.

Importance: 3/5

HF Daily Papers with >=100 upvotes bump applied.

Sources