Firehose

Filtered to tagged “embodied agents” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

30 JUL 2026 · Paper

This paper develops a framework, SpatialCLI, to help vision-language models (VLMs) better understand and use visual tools to make better decisions. By training VLMs to reason with spatial tools and then internalize those capabilities, SpatialCLI can improve the performance of VLMs in tasks that require visual reasoning.