🤖 AI Summary
This study addresses the scalability bottlenecks of LLM-as-a-Judge paradigms, including high computational costs, substantial latency, and opaque decision-making. To overcome these limitations, this work proposes programmatic distillation, pioneering the use of interpretable programs as substitutes for large language models (LLMs) in evaluation tasks. Specifically, we introduce PAJAMA, a system that aggregates judgments through a multi-program committee and incorporates a confidence-based routing mechanism to defer low-confidence samples back to LLMs. Experimental results demonstrate that PAJAMA achieves a 47-fold throughput improvement over 13B-parameter models while maintaining comparable accuracy. Furthermore, it reduces the cost of reward signal generation by two orders of magnitude. By preserving high precision, the proposed approach significantly enhances both evaluation efficiency and interpretability.
📝 Abstract
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, inference latency, and opaque decisions---limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at evaluation time, we distill its decision logic into a committee of programs that can score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and eight model families, we show that programmatic judges match the performance of a 13B-size LLM judge at 47x higher throughput. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs'verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.