21 upvotes · 30 JUL 2026 · Yang Zhou, Zixuan Huang, Sunzhu Li et al.
This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.
13 upvotes · 27 JUL 2026 · Zichao Lin, Yifeng Xie, Bowen Qu et al.
This paper introduces a benchmark to evaluate the atomic visual perception capabilities of large language models, which are often unable to accurately perceive visual information. Practitioners may care about this research because it provides a standardized way to measure and diagnose the limitations of visual perception in MLLMs.