TTPO: Test-Time Policy Optimization
This paper proposes a new method called Test-Time Policy Optimization (TTPO) that allows large language models to be trained without ground-truth labels, enabling test-time training. Practitioners may care about this because it enables models to be trained without labels, which can be difficult or expensive to obtain.