This paper develops a new method for image captioning that also grounds each phrase with a specific region of the image, allowing for more accurate and detailed descriptions. Practitioners might care about this work if they're building AI systems that need to understand and interact with the physical world.
Firehose
Filtered to Papers, tagged “panoptic segmentation” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives