OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
This paper develops a new generative model called OmniVAE that can jointly generate synchronized audio and video with fine-grained cross-modal correspondence, and its approach is expected to improve the quality of downstream text-to-audio-video generation tasks.