SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
This paper develops a system for generating speech and audio for various applications, including animation and video production, without reference recordings. Practitioners can use this system to create customized voices and control speaker styles, and to generate high-quality audio in complex scenarios.