Firehose

Filtered to tagged “KV Cache” · clear filters

All PeopleCompaniesPapersPodcastsHacker News

Browse by tag

3 AUG 2026 · Podcast · Latent Space: The AI Engineer Podcast

In this episode, Philip Kiely and Ali Taha from Baseten discuss the complexities and innovations in inference engineering for large AI models. They cover topics including model deployment, speculative decoding, quantization, hardware optimi…

29 APR 2026 · Podcast · Dwarkesh Podcast

Reiner Pope delivers a blackboard lecture on the mathematical and hardware principles behind training and serving large language models. He explains how batch size, sparsity, and various parallelism strategies (expert, pipeline) impact late…