This paper introduces Fathom, a technique to speed up decoding in large language models by selectively reading only the relevant parts of the key-value cache, reducing the computational cost and memory access. Practitioners in the field of natural language processing and deep learning may care about optimizing decoding efficiency for large models.
Firehose
Filtered to Papers, tagged “sparse decoding” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives