Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making
This paper introduces a new method, Multi-Head Latent Control, that allows large language models to make decisions at inference time, such as whether to use a stronger model or request additional information, without relying on costly and difficult-to-maintain external signals. Practitioners can use this method to improve the quality and efficiency of multi-model systems.