🤖 AI Summary
This study addresses the challenge that agent performance is constrained by model–framework compatibility, where training data cannot exhaustively cover the entire combinatorial space. To this end, we propose a joint routing method that leverages CP tensor decomposition to model such compatibility. By integrating decoupled representation learning with interaction term co-optimization, our approach accurately infers unobserved model–framework combinations from limited samples. Experimental results demonstrate that the proposed method surpasses baseline approaches in routing accuracy by 7.3% and achieves a substantial improvement of 15.8 percentage points on unobserved combinations. These findings indicate that our framework significantly enhances both the cost-effectiveness and generalization capability of agent selection.
📝 Abstract
Agent performance depends on both the underlying model and the harness that manages its tool use and execution. Selecting a suitable pair requires accounting for their compatibility, yet training samples may cover only a subset of the growing combination space. We introduce HM-Router, a routing method that jointly selects a model and harness for each query. It learns separate model and harness representations shared across routes, with an interaction term inspired by canonical polyadic (CP) tensor decomposition to capture how their compatibility varies with the query. This sharing allows training samples from observed pairs to inform predictions for unobserved combinations. We curate a benchmark from 12 public agent benchmarks, covering 293 routes, 73 models, and 25 harnesses. HM-Router exceeds the strongest evaluated learned baseline by 7.3 percentage points in mean routing accuracy and leads at all seven evaluated cost budgets on the six-benchmark subset. When 90% of routes have their training outcomes withheld, allowing unobserved combinations improves normalized accuracy by 15.8 points over restricting the same router to observed routes. HM-Router has also demonstrated training sample efficiency for new routes and components and generalization to unseen benchmarks. Our code and data are open-sourced at https://github.com/hmarkc/HM-Router.