Papers

Filtered to vision-language models · clear filter

Browse by term

continual learning 64reinforcement learning 34large language models 17benchmarking 12vision-language models 10generative models 8language models 8video generation 7multimodal models 6natural language processing 6robotics 6world models 6benchmarks 5diffusion models 5on-policy distillation 5policy optimization 5scalability 5self-distillation 5vision-language-action models 5autoregressive models 4computer vision 4diffusion transformers 4LLMs 4multimodal large language models 4verifiable rewards 4attention mechanisms 3embodied intelligence 3image editing 3long-term memory 3multimodal learning 3

Matching papers

HumanCLAW: Can Vision-Language Models Act Through a Body?

65 upvotes · 29 JUL 2026 · Siyao Li, Jiawei Gu, Shuai Liu et al.

This paper evaluates whether vision-language models can act through a physical body and how they can make decisions about what actions to take, without being hindered by issues like balance and motor control. Practitioners in AI and robotics might care because understanding how models interact with their physical bodies can help improve their ability to navigate and interact with the world.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

26 upvotes · 27 JUL 2026 · Senqiao Yang, Kaichen Zhang, Zhaoyang Jia et al.

This paper develops a new type of AI model that can understand and interact with both images and text in real-time, without needing a huge amount of training data. Practitioners may care about this model because it can be used in applications where fast and efficient visual perception is required.

Scaling Native Multimodal Pre-Training From Scratch

21 upvotes · 24 JUL 2026 · Haoyuan Wu, Aoqi Wu, Hai Wang et al.

This paper investigates how to scale large language models to also understand and interact with the physical world by training them on multiple types of data from scratch, allowing them to reason about both text and images. Practitioners might care because this could lead to more robust and versatile AI systems that can handle a wider range of tasks.

SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them

21 upvotes · 30 JUL 2026 · Yang Zhou, Zixuan Huang, Sunzhu Li et al.

This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.

Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment

14 upvotes · 15 JUL 2026 · Dwip Dalal, Shivansh Patel, Chahit Jain et al.

This paper proposes a new method for fine-tuning vision-language models on robot demonstrations to improve their performance on real-world tasks, by preventing the overwrite of pre-trained representations and aligning language and action predictions. Practitioners may care about this work because it aims to improve the generalizability and robustness of vision-language-action policies in real-world applications.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

13 upvotes · 28 JUL 2026 · Mingqiao Ye, Zhaochong An, Zhitong Gao et al.

This paper introduces Modus, a new type of any-to-any model that can predict any modality from any combination of others using only a decoder, without the need for specialized heads or losses. This approach can support a wide range of applications, such as chained generation and cross-modal self-verification, and has shown strong out-of-the-box performance on various benchmarks.

HPD-Parsing: Hierarchical Parallel Document Parsing

8 upvotes · 21 JUL 2026 · Shu Wei, Jingjing Wu, Lingshu Zhang et al.

This paper introduces HPD-Parsing, a new approach to document parsing that uses hierarchical parallel decoding to improve efficiency and throughput. Practitioners in natural language processing and computer vision might care because it could lead to faster and more accurate document parsing models.

Robostral Navigate

8 upvotes · 22 JUL 2026 · Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi et al.

This paper introduces Robostral Navigate, a vision-language model that enables robots to navigate using only a single monocular RGB camera, making it more scalable and cost-effective for deployment across various robotic platforms. Practitioners might care about this because it can simplify navigation tasks for robots in real-world environments.

SceneActBench: Can Agents Act on the 3D Scenes They See?

6 upvotes · 24 JUL 2026 · Yifei Zhao, Xiangxin Zhou, Wenhao Yang et al.

This paper introduces SceneActBench, a benchmark for evaluating the ability of vision-language model agents to perform actions in 3D scenes, and analyzes the performance of 11 different VLM configurations on this benchmark. Practitioners may care about this research because it helps to identify the limitations of current VLM agents in performing complex tasks in 3D environments.