MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
This paper introduces Modus, a new type of any-to-any model that can predict any modality from any combination of others using only a decoder, without the need for specialized heads or losses. This approach can support a wide range of applications, such as chained generation and cross-modal self-verification, and has shown strong out-of-the-box performance on various benchmarks.