OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
This paper proposes a method to compress large language models by allocating a fixed budget to select the most important tokens, allowing for efficient inference and reduced memory usage. Practitioners may care about this research if they work on optimizing the performance of large language models for real-world applications.