Papers

Filtered to vision-language models · clear filter

Browse by term

continual learning 83reinforcement learning 50large language models 12benchmarking 11benchmarks 11language models 11vision-language models 11robotics 7natural language processing 6world models 6generative models 5recursive self-improvement 5attention mechanisms 4diffusion Transformers 4multi-agent systems 4multimodal learning 4multimodal models 4on-policy distillation 4self-distillation 4self-supervised learning 4transformers 4video generation 4vision-language-action models 4agent-based systems 3agentic models 3agentic search 3autonomous systems 3coding agents 3diffusion models 3image generation 3

Matching papers

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

360 upvotes · 31 AUG 2026 · Xin Zhou, Zongchuang Zhao, Zhibo Yang et al.

This paper develops a vision-language model for autonomous driving that combines 3D perception, question answering, and motion planning. A practitioner might care because it demonstrates a promising approach to integrating multiple tasks in autonomous driving.

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

172 upvotes · 14 SEP 2026 · DeepCybo Team, Yu Bin, Haipeng Cao et al.

This paper develops a unified model that can understand physical environments, generate actions, and predict future states, using a combination of vision, language, and embodied interactions. Practitioners may care about this model because it could be used to create robots or other agents that can interact with and adapt to their physical surroundings.

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

155 upvotes · 18 AUG 2026 · Keyu Tu, Zhuowei Chen, Mengqi Huang et al.

This paper introduces a new benchmark for video generation tasks that require both achieving a desired outcome and maintaining semantic consistency with a reference image. Practitioners might care about this research because it can help evaluate and improve the performance of video generation models in real-world applications.

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

134 upvotes · 7 SEP 2026 · Soohyun Ryu, Sohee Kim, Eunho Yang

This paper helps large Vision-Language Models (LVLMs) better understand and reason about 3D scenes from 2D images, a key aspect of spatial intelligence, by training them on a synthetic dataset of block-stacking problems. Practitioners might care because improving spatial intelligence can lead to better performance on various visual tasks.

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

92 upvotes · 19 AUG 2026 · Yunhao Yang, Yuexin Bian, Yunjie Tian et al.

This paper introduces Co-RL, a framework for unsupervised multi-agent reinforcement learning that enables diverse and accurate reasoning in language and vision-language models. Practitioners can use Co-RL to improve their models' ability to reason and respond without relying on expensive ground-truth supervision.

Show-Harness: Just a VLM Agent Can Play Robots

89 upvotes · 9 SEP 2026 · Yanzhe Chen, Zechen Bai, Zhijun Cao et al.

This paper shows how a vision-language model (VLM) can control robots without needing extensive pretraining or specialized hardware, by providing a compact interface that links the model's intentions to specific actions. Practitioners might care about this because it could make robots more accessible and user-friendly.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

68 upvotes · 30 JUL 2026 · Qiushi Sun, Kanzhi Cheng, Yian Wang et al.

This paper develops a standardized evaluation method for computer-using agents (CUAs) to ensure they fulfill task instructions, using vision-language models (VLMs) as judges. Practitioners can benefit from this work by using reliable and cost-effective reward signals for CUA training.

HumanCLAW: Can Vision-Language Models Act Through a Body?

65 upvotes · 29 JUL 2026 · Siyao Li, Jiawei Gu, Shuai Liu et al.

This paper evaluates whether vision-language models can act through a physical body and how they can make decisions about what actions to take, without being hindered by issues like balance and motor control. Practitioners in AI and robotics might care because understanding how models interact with their physical bodies can help improve their ability to navigate and interact with the world.

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

64 upvotes · 3 SEP 2026 · Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham et al.

This paper proposes a new approach to Large Language Models (LLMs) that uses diffusion to speed up generation without sacrificing quality, allowing for faster and more efficient language processing. Practitioners may care about this research if they need to process large amounts of text quickly and efficiently.

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

52 upvotes · 21 AUG 2026 · Yibo Hu, Yu Qian, Mao Gu et al.

This paper proposes a model called TLive-Omni that can understand and process multiple types of data (image, video, audio, text) from e-commerce live streaming. A practitioner might care about this model because it can be used to analyze and respond to live-commerce streams in a more accurate and efficient way.

Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

46 upvotes · 18 AUG 2026 · Hongyan Feng, Sunlai Chen, Xuanyu Liu et al.

This paper proposes a new framework for embodied navigation that overcomes limitations of existing methods by reformulating navigation into a 2D visual space, introducing selective reasoning and memory mechanisms, and designing an efficient alignment paradigm. Practitioners caring about efficient navigation for AI agents might care about this research.