Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
This paper introduces Chimera, a hybrid visual diffusion transformer that efficiently processes text, image, and video tokens to generate high-resolution images, videos, and multimodal context. Practitioners might care about this paper because it provides a scalable solution for large-scale visual generation tasks.