This paper proposes a new framework for joint multimodal representation learning and generation, allowing for flexible-length aligned transmodal tokens that can be used for both retrieval and generation tasks. Practitioners might care about this paper because it shows how to improve generative performance by training a shared multimodal encoder alongside downstream models.
Firehose
Filtered to Papers, tagged “image captioning” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives