158 upvotes · 14 AUG 2026 · Yijiang Li, Yijun Liang, Yunjie Tian et al.
This paper proposes a new method called Self-Supervised Visual On-Policy Distillation (S^2VOPD) that generates learning signals from asymmetric augmented views of images, allowing for effective on-policy learning without privileged information. Practitioners might care about this paper because it presents a simple yet effective way to improve performance on various perception benchmarks.
78 upvotes · 8 SEP 2026 · Youngrok Park, Sangmin Bae, Hojung Jung et al.
This paper introduces a method called On-Policy Reverse Distillation that helps stronger models learn from weaker supervisors by selectively amplifying the parts of the weaker model's guidance that are most useful to the stronger model. This can lead to faster and more efficient learning, especially in situations where it's expensive to retrain models from scratch.