65 upvotes · 29 JUL 2026 · Siyao Li, Jiawei Gu, Shuai Liu et al.
This paper evaluates whether vision-language models can act through a physical body and how they can make decisions about what actions to take, without being hindered by issues like balance and motor control. Practitioners in AI and robotics might care because understanding how models interact with their physical bodies can help improve their ability to navigate and interact with the world.
26 upvotes · 27 JUL 2026 · Senqiao Yang, Kaichen Zhang, Zhaoyang Jia et al.
This paper develops a new type of AI model that can understand and interact with both images and text in real-time, without needing a huge amount of training data. Practitioners may care about this model because it can be used in applications where fast and efficient visual perception is required.
21 upvotes · 24 JUL 2026 · Haoyuan Wu, Aoqi Wu, Hai Wang et al.
This paper investigates how to scale large language models to also understand and interact with the physical world by training them on multiple types of data from scratch, allowing them to reason about both text and images. Practitioners might care because this could lead to more robust and versatile AI systems that can handle a wider range of tasks.
21 upvotes · 30 JUL 2026 · Yang Zhou, Zixuan Huang, Sunzhu Li et al.
This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.
14 upvotes · 15 JUL 2026 · Dwip Dalal, Shivansh Patel, Chahit Jain et al.
This paper proposes a new method for fine-tuning vision-language models on robot demonstrations to improve their performance on real-world tasks, by preventing the overwrite of pre-trained representations and aligning language and action predictions. Practitioners may care about this work because it aims to improve the generalizability and robustness of vision-language-action policies in real-world applications.
13 upvotes · 28 JUL 2026 · Mingqiao Ye, Zhaochong An, Zhitong Gao et al.
This paper introduces Modus, a new type of any-to-any model that can predict any modality from any combination of others using only a decoder, without the need for specialized heads or losses. This approach can support a wide range of applications, such as chained generation and cross-modal self-verification, and has shown strong out-of-the-box performance on various benchmarks.
8 upvotes · 21 JUL 2026 · Shu Wei, Jingjing Wu, Lingshu Zhang et al.
This paper introduces HPD-Parsing, a new approach to document parsing that uses hierarchical parallel decoding to improve efficiency and throughput. Practitioners in natural language processing and computer vision might care because it could lead to faster and more accurate document parsing models.
8 upvotes · 22 JUL 2026 · Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi et al.
This paper introduces Robostral Navigate, a vision-language model that enables robots to navigate using only a single monocular RGB camera, making it more scalable and cost-effective for deployment across various robotic platforms. Practitioners might care about this because it can simplify navigation tasks for robots in real-world environments.
6 upvotes · 24 JUL 2026 · Yifei Zhao, Xiangxin Zhou, Wenhao Yang et al.
This paper introduces SceneActBench, a benchmark for evaluating the ability of vision-language model agents to perform actions in 3D scenes, and analyzes the performance of 11 different VLM configurations on this benchmark. Practitioners may care about this research because it helps to identify the limitations of current VLM agents in performing complex tasks in 3D environments.
5 upvotes · 30 JUL 2026 · Yao Xiao, Reuben Tan, Zhen Zhu et al.
This paper proposes a new approach to improve vision-language models for visual retrieval, which can handle long visual contexts and large numbers of distractors. Practitioners might care because it can lead to better performance on image and video benchmarks.