PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
This paper introduces a benchmark to evaluate the atomic visual perception capabilities of large language models, which are often unable to accurately perceive visual information. Practitioners may care about this research because it provides a standardized way to measure and diagnose the limitations of visual perception in MLLMs.