72 upvotes · 27 AUG 2026 · Aozhe Wang, Zhengxi Lu, Jianze Wang et al.
This paper proposes a new method called Test-Time Policy Optimization (TTPO) that allows large language models to be trained without ground-truth labels, enabling test-time training. Practitioners may care about this because it enables models to be trained without labels, which can be difficult or expensive to obtain.
47 upvotes · 25 AUG 2026 · Wei Zhou, Xiongwei Zhu, Lingdong Kong et al.
This paper introduces a new method for improving diffusion models by using self-distillation to align them with human preferences and task-specific objectives. Practitioners might care about this approach because it can lead to more efficient and analyzable alignment of diffusion models with human goals.