This paper presents a new technique to reduce memory usage and speed up inference for large neural networks, allowing them to run on consumer hardware with limited memory. A practitioner might care about this because it enables the deployment of large models in edge devices and reduces the need for expensive storage.
Firehose
Filtered to Papers, tagged “deep learning inference” · clear filters
Browse: People · Companies · Papers · Podcasts · Hacker News · Deep dives