In this episode, Philip Kiely and Ali Taha from Baseten discuss the complexities and innovations in inference engineering for large AI models. They cover topics including model deployment, speculative decoding, quantization, hardware optimi…
Inference engineeringSpeculative decodingQuantizationModel deploymentTool callingKV cacheTensor parallelismExpert parallelismGPU hardwareRubin GPUVideo diffusionAutoregressive modelsDiffusion modelsTraining-inference convergenceContinual learningOpen source modelsInference infrastructureModel optimizationLatency vs throughputMulti-modal models