Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in large language model (LLM) routing where escalation signals, such as semantic entropy, frequently exhibit spurious correlations due to a "difficulty baseline trap," thereby misleading system decisions. To overcome this, we propose a general evaluation framework comprising five diagnostic checks, integrated with AUROC analysis and cost-benefit modeling, to rigorously distinguish genuine performance gains from spurious correlations and predict the failure boundaries of caching strategies. Evaluated on the GSM8K benchmark, the semantic entropy signal achieves an AUROC of 0.871 and improves routing accuracy by 9%, validating the effectiveness of the proposed approach. This work establishes a systematic paradigm for assessing the reliability of routing signals in LLM systems.
📝 Abstract
Deciding when to escalate a query from a small language model to a larger one requires a cheap signal that predicts, before the large model is called, whether escalating would help. Semantic entropy, originally developed to detect hallucinations, is a natural candidate: it measures how much a model's sampled answers disagree in meaning, and high disagreement often signals an unreliable answer. We test it across three benchmarks and two model families. On GSM8K, with a small/large pair about twelve times apart in size, semantic entropy reliably distinguishes the small model's mistakes (AUROC 0.871) and improves routed accuracy over random escalation by up to nine points at matched cost. An earlier strong-looking result on a synthetic benchmark proved misleading: a simple rule based only on question difficulty, with no model involved, matched semantic entropy almost exactly. This paper's main contribution is a set of checks that catch this before it is reported as real. We show that scoring a cheap, question-only difficulty estimate alongside any signal reveals whether the signal adds real information or just tracks how hard a question looks; that two reasonable definitions of "escalation worked" can produce very different results on the same data; that a benchmark can leave almost no room for any signal to beat simply always using the large model; and that the true cost of live sampling can make routing more expensive than calling the large model directly. For a cheaper alternative that reuses cached past outcomes, we show how to predict whether it will work on a new dataset -- confirmed by correctly forecasting a collapse from AUROC 0.908 to chance level (0.518) ahead of time. We offer these as a general checklist for evaluating escalation signals.
Problem

Research questions and friction points this paper is trying to address.

LLM routing
escalation signals
semantic entropy
evaluation pitfalls
query escalation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM Routing
Semantic Entropy
Escalation Signals
Evaluation Checklist
Question Difficulty Baseline
🔎 Similar Papers
No similar papers found.
R
Ramin Pishehvar
Cisco Systems, Inc., San Jose, CA, USA
A
Andrea Morandi
Cisco Systems, Inc., San Jose, CA, USA
Mahesh Viswanathan
Mahesh Viswanathan
University of Illinois, Urbana-Champaign
formal verificationautomata theoryconcurrency theorylogic