Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
This paper introduces Grouped Value Attention (GVA), a technique to reduce the memory footprint of Transformer decoding by storing grouped values and reconstructing content keys. Practitioners might care about this because it could lead to faster and more efficient models.