360 upvotes · 31 AUG 2026 · Xin Zhou, Zongchuang Zhao, Zhibo Yang et al.
This paper develops a vision-language model for autonomous driving that combines 3D perception, question answering, and motion planning. A practitioner might care because it demonstrates a promising approach to integrating multiple tasks in autonomous driving.
172 upvotes · 14 SEP 2026 · DeepCybo Team, Yu Bin, Haipeng Cao et al.
This paper develops a unified model that can understand physical environments, generate actions, and predict future states, using a combination of vision, language, and embodied interactions. Practitioners may care about this model because it could be used to create robots or other agents that can interact with and adapt to their physical surroundings.
155 upvotes · 18 AUG 2026 · Keyu Tu, Zhuowei Chen, Mengqi Huang et al.
This paper introduces a new benchmark for video generation tasks that require both achieving a desired outcome and maintaining semantic consistency with a reference image. Practitioners might care about this research because it can help evaluate and improve the performance of video generation models in real-world applications.
134 upvotes · 7 SEP 2026 · Soohyun Ryu, Sohee Kim, Eunho Yang
This paper helps large Vision-Language Models (LVLMs) better understand and reason about 3D scenes from 2D images, a key aspect of spatial intelligence, by training them on a synthetic dataset of block-stacking problems. Practitioners might care because improving spatial intelligence can lead to better performance on various visual tasks.
92 upvotes · 19 AUG 2026 · Yunhao Yang, Yuexin Bian, Yunjie Tian et al.
This paper introduces Co-RL, a framework for unsupervised multi-agent reinforcement learning that enables diverse and accurate reasoning in language and vision-language models. Practitioners can use Co-RL to improve their models' ability to reason and respond without relying on expensive ground-truth supervision.
89 upvotes · 9 SEP 2026 · Yanzhe Chen, Zechen Bai, Zhijun Cao et al.
This paper shows how a vision-language model (VLM) can control robots without needing extensive pretraining or specialized hardware, by providing a compact interface that links the model's intentions to specific actions. Practitioners might care about this because it could make robots more accessible and user-friendly.
68 upvotes · 30 JUL 2026 · Qiushi Sun, Kanzhi Cheng, Yian Wang et al.
This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.
65 upvotes · 29 JUL 2026 · Siyao Li, Jiawei Gu, Shuai Liu et al.
This paper evaluates whether vision-language models can act through a physical body and how they can make decisions about what actions to take, without being hindered by issues like balance and motor control. Practitioners in AI and robotics might care because understanding how models interact with their physical bodies can help improve their ability to navigate and interact with the world.
64 upvotes · 3 SEP 2026 · Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham et al.
This paper proposes a new approach to Large Language Models (LLMs) that uses diffusion to speed up generation without sacrificing quality, allowing for faster and more efficient language processing. Practitioners may care about this research if they need to process large amounts of text quickly and efficiently.
52 upvotes · 21 AUG 2026 · Yibo Hu, Yu Qian, Mao Gu et al.
This paper proposes a model called TLive-Omni that can understand and process multiple types of data (image, video, audio, text) from e-commerce live streaming. A practitioner might care about this model because it can be used to analyze and respond to live-commerce streams in a more accurate and efficient way.
46 upvotes · 18 AUG 2026 · Hongyan Feng, Sunlai Chen, Xuanyu Liu et al.
This paper proposes a new framework for embodied navigation that overcomes limitations of existing methods by reformulating navigation into a 2D visual space, introducing selective reasoning and memory mechanisms, and designing an efficient alignment paradigm. Practitioners caring about efficient navigation for AI agents might care about this research.