Enhancing Rubric-based RL via Self-Distillation
This paper improves a type of reinforcement learning (RL) called rubric-based RL, which helps large language models (LLMs) perform well on open-ended tasks. A practitioner might care about this paper because it addresses a common problem in RL, where some criteria (or rules) are not explored properly, and it shows that its new method can improve performance on these tasks.