Firehose

Filtered to tagged “sparse decoding” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

15 SEP 2026 · Paper

This paper introduces Fathom, a technique to speed up decoding in large language models by selectively reading only the relevant parts of the key-value cache, reducing the computational cost and memory access. Practitioners in the field of natural language processing and deep learning may care about optimizing decoding efficiency for large models.