Firehose

Filtered to tagged “multimodal large language models” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

14 SEP 2026 · Paper

This paper introduces HarnessVLN, a training-free framework for embodied navigation that uses a unified tool interface to validate proposed actions against spatial evidence and task progress, allowing agents to generalize and learn from multimodal large language models.

13 SEP 2026 · Paper

This paper introduces OmniHarness, a framework that enables generalizable visual generation by learning symbolic policies that can be applied to multiple tasks, allowing for more efficient and effective visual generation. Practitioners might care about this research because it could lead to more robust and adaptable visual generation systems.