Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
This paper improves reinforcement learning for post-training language model reasoners by addressing two problems in existing methods: identical advantages for distinct reward profiles and fixed relative weights for all objectives. A new method, SA-MRPO, dynamically reallocates optimization effort toward under-optimized objectives while maintaining performance on well-satisfied objectives.