🤖 AI Summary
This work addresses the inefficiency of current AI programming systems that uniformly employ high-end large language models for all tasks, incurring excessive inference costs for routine software engineering activities. The authors propose a dynamic routing mechanism that leverages code health metrics and task metadata to assign tasks to the lowest-cost model tier—lightweight, standard, or heavyweight—that meets required quality thresholds. Introducing code quality signals into model selection for the first time, they formulate verifiable cost-effectiveness conditions and establish a rigorous evaluation protocol to quantify the impact of individual code health subfactors on routing decisions. Evaluated on SWE-bench Lite using a three-tier strategy—comprising heuristic thresholds, machine learning classifiers, and an ideal oracle—the approach demonstrates significant cost reduction without compromising output quality when lightweight models achieve pass rates on high-health code exceeding their relative cost advantage and the code health effect size reaches 𝑝̂ ≥ 0.56.
📝 Abstract
Context: AI coding agents route every task to a single frontier large language model (LLM), paying premium inference cost even when many tasks are routine.
Objectives: We propose Triage, a framework that uses code health metrics -- indicators of software maintainability -- as a routing signal to assign each task to the cheapest model tier whose output passes the same verification gate as the expensive model.
Methods: Triage defines three capability tiers (light, standard, heavy -- mirroring, e.g., Haiku, Sonnet, Opus) and routes tasks based on pre-computed code health sub-factors and task metadata. We design an evaluation comparing three routing policies on SWE-bench Lite (300 tasks across three model tiers): heuristic thresholds, a trained ML classifier, and a perfect-hindsight oracle.
Results: We analytically derived two falsifiable conditions under which the tier-dependent asymmetry (medium LLMs benefit from clean code while frontier models do not) yields cost-effective routing: the light-tier pass rate on healthy code must exceed the inter-tier cost ratio, and code health must discriminate the required model tier with at least a small effect size ($\hat{p} \geq 0.56$).
Conclusion: Triage transforms a diagnostic code quality metric into an actionable model-selection signal. We present a rigorous evaluation protocol to test the cost--quality trade-off and identify which code health sub-factors drive routing decisions.