DAPD: Dual-Anchored Policy Distillation
Identifies a 'privilege illusion' failure mode in on-policy self-distillation, where a student model learns to rely on privileged teacher-time information it cannot access at inference. DAPD fixes this with dual-path anchoring and dual-source anchoring, beating prior on-policy self-distillation by +2.00 points averaged over six benchmarks on Qwen3-4B.
Why it matters
HuggingFace Daily Papers entry with 73 upvotes; names a concrete, previously under-described failure mode in self-distillation pipelines increasingly used to compress reasoning models.
Importance: 3/5
Notable paper (73 HF Daily Papers upvotes, below the +1 bump threshold).
Sources
official
HuggingFace Daily Papers