OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
This paper introduces a new benchmark (OVEarth-Bench) to evaluate open-vocabulary Earth observation models, focusing on both category breadth and query diversity. Practitioners in this field can benefit from understanding the importance of developing more realistic and diverse benchmarks for reliable model evaluation.