DAPD: Dual-Anchored Policy Distillation
This paper addresses a problem in self-distillation, where a student model learns to mimic the behavior of a privileged teacher, but performs poorly at inference due to a "privilege illusion". The authors propose a new method, Dual-Anchored Policy Distillation, to resolve this issue and improve performance.