DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
This paper introduces a systematic way to construct pretraining mixtures for Vision Language Models (VLMs) by breaking down the process into two parts: deciding which classes to combine and how to allocate data within each class. Practitioners can use this approach to improve the quality and diversity of their VLMs.