This paper introduces Chimera, a hybrid visual diffusion transformer that efficiently processes text, image, and video tokens to generate high-resolution images, videos, and multimodal context. Practitioners might care about this paper because it provides a scalable solution for large-scale visual generation tasks.
Firehose
Filtered to Papers, tagged “multimodal generation” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives