SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
This paper proposes a new method to sparsify attention in Transformers, called Simple Attention Sparsification (SAS), which optimizes context ranking end-to-end with the language modeling loss. Practitioners might care because SAS can improve performance on downstream tasks by using attention budgets more effectively.