#distillation
- S²VOPD: self-supervised visual on-policy distillation lifts Qwen3.5-4B from 70.7% to 77.4% on six fine-grained perception benchmarks UC San Diego research
- DOPD: Dual On-Policy Distillation with Advantage-Aware Token Routing research
- NSA, CISA and FBI accuse six Chinese AI firms of industrial-scale distillation of US models industry
- Anthropic threat report names Moonshot, DeepSeek and Alibaba in distillation campaigns Anthropic research
- Moebius: 0.2B Lightweight Image Inpainting Framework Matches 11.9B FLUX Model Huazhong University of Science and Technology research
- NeoHorse-1: recursive self-improvement via agentic post-training TokenRhythm research
- Causal Forcing++: 2-Step Distillation Enables Real-Time Interactive Video Generation Tsinghua University research
- SDAR: Self-Distilled Agentic Reinforcement Learning for Multi-Turn Agents Zhejiang University / Meituan research
- ThoughtFold: Introspective Preference Learning Cuts Reasoning Tokens by 56% Without Accuracy Loss research
- DAPD: Dual-Anchored Policy Distillation research
- On-Policy Self-Distillation without Any Supervision research
- TTPO: Test-Time Policy Optimization Zhejiang University research
- Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher research
- AnyFlow: Any-Step Video Diffusion with On-Policy Flow Map Distillation MIT / NVIDIA research
- TrOPD: Trust-Region On-Policy Distillation Stabilizes LLM Training When Teacher-Student Gap Is Large Samsung Research research
- DanceOPD: On-Policy Generative Field Distillation for Unified Image Generation ByteDance Seed research
- Pass the Baton: Trajectory-Relayed On-Policy Distillation Zhejiang University research
- ReWorld: An Interactive World Model with Long-Horizon Memory TongyiLab (Alibaba) research
- Anthropic Accuses Alibaba of Largest Known Claude Distillation Attack: 28.8M Conversations Anthropic industry
- White House official accuses Moonshot AI of distilling Anthropic's Fable to build Kimi K3; experts push back Moonshot AI industry
- On the Geometry of On-Policy Distillation: A Training Paradigm Distinct from SFT and RLVR Hong Kong University of Science and Technology research
- Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Rutgers University research
- ZPPO: Teacher-in-Prompts Knowledge Distillation Outperforms Gradient Methods for Small Reasoners NVIDIA research
- Weak-to-Strong Generalization via Direct On-Policy Distillation ByteDance / Tsinghua University research
- Weak-to-Strong On-Policy Distillation Microsoft Research / University of Maryland / MBZUAI research
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory research
- TRL v1.11.0 rewrites trl vllm-serve (~10x smaller) and ships AsyncDistillationTrainer Hugging Face tools
- Does on-policy distillation really distill? Teacher-free OPSA beats it on AIME24 Purdue University research
- Rethinking On-Policy Distillation of LLMs II: near-full gains from a single training example research
- OPRD: eliciting weak-to-strong generalization with on-policy reverse distillation KAIST AI research
- When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation research