SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
This paper proposes a method to combine reinforcement learning with verifiable rewards and on-policy distillation to improve performance on complex tasks, and shows that this method can lead to more stable training and better results.