48 upvotes · 30 JUL 2026 · Qixun Wang, Yang Shi, Letian Cheng et al.
This paper proposes a new approach to agentic visual reasoning, which helps large language models (LLMs) perform better on complex tasks by using tools more efficiently. Practitioners might care about this research because it aims to improve the performance of LLMs on challenging problems.
25 upvotes · 17 JUL 2026 · Jiarui Zhang, Muzi Tao, Shangshang Wang et al.
This paper introduces a benchmark called ActiveVision to measure whether large language models (LLMs) exercise active observation, and finds that current LLMs are not robust in this regard, performing poorly on tasks that require repeated visual perception.
13 upvotes · 27 JUL 2026 · Zichao Lin, Yifeng Xie, Bowen Qu et al.
This paper introduces a benchmark to evaluate the atomic visual perception capabilities of large language models, which are often unable to accurately perceive visual information. Practitioners may care about this research because it provides a standardized way to measure and diagnose the limitations of visual perception in MLLMs.
11 upvotes · 31 JUL 2026 · Yingmao Miao, Pengfei Zhang, Xiaochen Lv et al.
This paper develops a new method to evaluate and verify image editing consistency across multiple references, addressing a challenge in reinforcement learning for multi-reference editing. Practitioners may care about this approach as it enables more accurate and reliable reinforcement learning for image editing tasks.