This paper investigates how text conditioning affects visual generation and proposes ways to improve it, leading to better performance on various benchmarks. Practitioners might care about the findings to develop more effective text-to-image models.
Firehose
Filtered to Papers, tagged “Diffusion models” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives