This paper introduces Chimera, a hybrid visual diffusion transformer that efficiently processes text, image, and video tokens to generate high-resolution images, videos, and multimodal context. Practitioners might care about this paper because it provides a scalable solution for large-scale visual generation tasks.
Firehose
Filtered to Papers, tagged “sparse mixture-of-experts” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives
Browse by tag
This paper investigates how sparse mixture-of-experts language models route tokens to multiple experts and how this routing affects their performance. Practitioners may care because understanding how to optimize these models can lead to better language understanding and generation.