🤖 AI Summary
This study addresses the limitations of heuristic-dependent multi-LLM team formation, which lacks computable complementarity metrics and consequently constrains performance. To overcome this, we propose a heterogeneity-based team selection framework that introduces the first computable and interpretable complementarity measurement system. By quantifying model capabilities, error decorrelation, and predictive behavioral diversity through offline profiling, the framework constructs a joint quality-complementarity optimization objective and employs greedy search to efficiently identify optimal teams. This approach transforms team formation from empirical heuristics into a standardized combinatorial optimization problem. Extensive evaluations across multiple benchmarks demonstrate that our method consistently outperforms quality-only baselines, establishing a reusable selection paradigm for multi-LLM systems.
📝 Abstract
The performance ceiling of an LLM team is constrained not only by individual model capabilities, but also by inter-member error resonance and predictive differences. Although heterogeneous teaming is often observed to be effective in practice, existing approaches lack complementarity metrics that are computable, interpretable, and optimizable, leaving team composition to rely on heuristics. We propose a heterogeneity-driven team selection framework that performs offline profiling to characterize individual capability along with two complementary signals: one captures decorrelation in error patterns to reduce co-failures, while the other measures divergence in predictive behavior to capture strategy diversity. We formulate team selection as a standardized quality--complementarity combinatorial objective and apply an efficient greedy search to select a small team from a candidate pool. Experiments across multiple benchmarks demonstrate that our framework consistently outperforms quality-only baselines under controlled candidate pools and team sizes, establishing reusable selection principles for multi-LLM systems.