Triage: Routing Software Engineering Tasks to Cost-Effective LLM Tiers via Code Quality Signals

📅 2026-04-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inefficiency of current AI programming systems that uniformly employ high-end large language models for all tasks, incurring excessive inference costs for routine software engineering activities. The authors propose a dynamic routing mechanism that leverages code health metrics and task metadata to assign tasks to the lowest-cost model tier—lightweight, standard, or heavyweight—that meets required quality thresholds. Introducing code quality signals into model selection for the first time, they formulate verifiable cost-effectiveness conditions and establish a rigorous evaluation protocol to quantify the impact of individual code health subfactors on routing decisions. Evaluated on SWE-bench Lite using a three-tier strategy—comprising heuristic thresholds, machine learning classifiers, and an ideal oracle—the approach demonstrates significant cost reduction without compromising output quality when lightweight models achieve pass rates on high-health code exceeding their relative cost advantage and the code health effect size reaches 𝑝̂ ≥ 0.56.

Technology Category

Planning, Routing, and Scheduling: Planning with Language ModelsNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageSearch and Optimization: Sampling/Simulation-based Search

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
Context: AI coding agents route every task to a single frontier large language model (LLM), paying premium inference cost even when many tasks are routine. Objectives: We propose Triage, a framework that uses code health metrics -- indicators of software maintainability -- as a routing signal to assign each task to the cheapest model tier whose output passes the same verification gate as the expensive model. Methods: Triage defines three capability tiers (light, standard, heavy -- mirroring, e.g., Haiku, Sonnet, Opus) and routes tasks based on pre-computed code health sub-factors and task metadata. We design an evaluation comparing three routing policies on SWE-bench Lite (300 tasks across three model tiers): heuristic thresholds, a trained ML classifier, and a perfect-hindsight oracle. Results: We analytically derived two falsifiable conditions under which the tier-dependent asymmetry (medium LLMs benefit from clean code while frontier models do not) yields cost-effective routing: the light-tier pass rate on healthy code must exceed the inter-tier cost ratio, and code health must discriminate the required model tier with at least a small effect size ($\hat{p} \geq 0.56$). Conclusion: Triage transforms a diagnostic code quality metric into an actionable model-selection signal. We present a rigorous evaluation protocol to test the cost--quality trade-off and identify which code health sub-factors drive routing decisions.
Problem

Research questions and friction points this paper is trying to address.

LLM cost efficiency
software engineering tasks
code quality
model tier routing
inference cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

code health
LLM tiering
cost-effective routing
model selection
software maintainability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lech Madeyski
Wroclaw University of Science and Technology, Poland