Firehose

Filtered to Papers, tagged “long-context models” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

13 SEP 2026 · Paper

This paper addresses the issue of memory peak allocation in Mixture-of-Experts (MoE) models during long-context training, proposing four different techniques to reduce memory usage without compromising performance. Practitioners might care about these techniques to train larger MoE models with longer context lengths.