Rethinking On-Policy Distillation of Large Language Models II: One Training Example
This paper investigates the role of training data in on-policy distillation, a technique used to improve large language models, and finds that even a single query can lead to significant improvements, but the process is slow and algorithm-starved.