This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.
Firehose
Filtered to Papers, tagged “embodied agents” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives