Codifying the Judge: Scalable Evaluation via Program Distillation
This paper proposes a way to make automated evaluation systems more efficient, transparent, and reliable by distilling the decision logic of large language models into smaller, programmatic judges that can be easily inspected and edited. Practitioners might care because this approach could help reduce costs and improve the scalability of automated evaluation systems.