This paper addresses the issue of memory peak allocation in Mixture-of-Experts (MoE) models during long-context training, proposing four different techniques to reduce memory usage without compromising performance. Practitioners might care about these techniques to train larger MoE models with longer context lengths.
Firehose
Filtered to Papers, tagged “activation management” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives