29 upvotes · 27 JUL 2026 · Yifan Ye, Yankai Fu, Yaoxu Lv et al.
This paper proposes a hierarchical structure for organizing data sources for embodied manipulation, aiming to balance scalability and robot alignment. Practitioners might care about this work if they're building or deploying embodied agents that require diverse and high-quality data.
5 upvotes · 27 JUL 2026 · Jiahao Xie, Zhongbin Guo, Qianle Wang et al.
This paper introduces a systematic way to construct pretraining mixtures for Vision Language Models (VLMs) by breaking down the process into two parts: deciding which classes to combine and how to allocate data within each class. Practitioners can use this approach to improve the quality and diversity of their VLMs.