The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Inference engineeringSpeculative decodingQuantizationModel deploymentTool callingKV cacheTensor parallelismExpert parallelismGPU hardwareRubin GPUVideo diffusionAutoregressive modelsDiffusion modelsTraining-inference convergenceContinual learningOpen source modelsInference infrastructureModel optimizationLatency vs throughputMulti-modal models
In this episode, Philip Kiely and Ali Taha from Baseten discuss the complexities and innovations in inference engineering for large AI models. They cover topics including model deployment, speculative decoding, quantization, hardware optimization, and emerging trends in video and audio model inference. The conversation also explores the convergence of training and inference, and the future of AI infrastructure with new hardware like NVIDIA's Rubin GPU.