This paper develops a unified model that can understand physical environments, generate actions, and predict future states, using a combination of vision, language, and embodied interactions. Practitioners may care about this model because it could be used to create robots or other agents that can interact with and adapt to their physical surroundings.
Firehose
Filtered to Papers, tagged “multimodal learning” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives