Firehose

Filtered to Papers, tagged “representation” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

30 JUL 2026 · Paper

This paper investigates how sparse mixture-of-experts language models route tokens to multiple experts and how this routing affects their performance. Practitioners may care because understanding how to optimize these models can lead to better language understanding and generation.