DAPD: Dual-Anchored Policy Distillation

Research official 1 src. ~1 min

Identifies a 'privilege illusion' failure mode in on-policy self-distillation, where a student model learns to rely on privileged teacher-time information it cannot access at inference. DAPD fixes this with dual-path anchoring and dual-source anchoring, beating prior on-policy self-distillation by +2.00 points averaged over six benchmarks on Qwen3-4B.

Why it matters

HuggingFace Daily Papers entry with 73 upvotes; names a concrete, previously under-described failure mode in self-distillation pipelines increasingly used to compress reasoning models.

Importance: 3/5

Notable paper (73 HF Daily Papers upvotes, below the +1 bump threshold).

Sources