256 upvotes · 29 JUL 2026 · Zeyu Zhang, Ziliang Guo, Yihang Sun et al.
This paper introduces a new type of AI model called memory foundation models, which allows the model to learn and retain information internally, rather than relying on external memory modules. This could be useful for practitioners who want to build more efficient and flexible AI agents.
29 upvotes · 27 JUL 2026 · Yifan Ye, Yankai Fu, Yaoxu Lv et al.
This paper proposes a hierarchical structure for organizing data sources for embodied manipulation, aiming to balance scalability and robot alignment. Practitioners might care about this work if they're building or deploying embodied agents that require diverse and high-quality data.
23 upvotes · 29 JUL 2026 · Siyu Yan, Zhuoran Yan, Haiying Xu et al.
This paper evaluates how multimodal large language models use intermediate visual states during reasoning and finds that these visual states are not as crucial as previously thought, but can still impact model performance under certain conditions. Practitioners might care because understanding how these models use visual states can help improve their performance and reliability.
13 upvotes · 28 JUL 2026 · Mingqiao Ye, Zhaochong An, Zhitong Gao et al.
This paper introduces Modus, a new type of any-to-any model that can predict any modality from any combination of others using only a decoder, without the need for specialized heads or losses. This approach can support a wide range of applications, such as chained generation and cross-modal self-verification, and has shown strong out-of-the-box performance on various benchmarks.
5 upvotes · 22 JUL 2026 · Md Tanvirul Alam
This paper introduces a new environment called Trace, which allows vision-language models to reason across multiple domains and tasks, using a taxonomy-guided approach. Practitioners might care because this work could lead to more generalizable and transferable AI models.
5 upvotes · 27 JUL 2026 · Jiahao Xie, Zhongbin Guo, Qianle Wang et al.
This paper introduces a systematic way to construct pretraining mixtures for Vision Language Models (VLMs) by breaking down the process into two parts: deciding which classes to combine and how to allocate data within each class. Practitioners can use this approach to improve the quality and diversity of their VLMs.