You're Hired: Strategic Model Selection for LLM Collaboration

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that multi-LLM collaborative systems are constrained by predefined model pools and lack efficient selection mechanisms. It proposes a taxonomy of nine model selection algorithms encompassing descriptive diversity, capability-aware behavioral diversity, and LLM recruiter strategies, which are systematically evaluated within a multi-agent framework across mathematical, coding, question-answering, and reasoning tasks. Experimental results demonstrate that capability- and training-based strategies significantly reduce selection variance, with the optimal algorithm achieving up to 36.1% performance improvement over baselines. Furthermore, the proposed approach effectively filters out unsafe models while exhibiting strong generalization on out-of-distribution tasks.
📝 Abstract
While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systematically evaluate a taxonomy of 9 selection algorithms ranging from diversity of model descriptions, capability-aware behavioral diversity, and LLM-based recruiters. We conduct extensive experiments across two candidate pools of 10 and 32 models, deployed in four model collaboration algorithms, and evaluated across tasks spanning math, coding, QA, and reasoning. Results demonstrate that successful selection algorithms greatly outperform random or heuristics-based teams such as merely selecting the models with top individual performance, by up to 36.1% across settings. Specifically, capability- and training-based selection strategies alleviate selection variance and achieve the best performance, which we recommend to employ before deploying real-world multi-LLM systems. Further analysis reveals that larger candidate pools pose greater challenges to shallow selection heuristics, while algorithms grounded in interacting with candidate models and understanding model capability robustly filter out misaligned, unsafe models, as well as generalizing to novel, out-of-distribution tasks. Together, we establish that principled and informed team selection is critical and present strong model selection algorithms for assembling effective multi-LLM systems.
Problem

Research questions and friction points this paper is trying to address.

Model Selection
Multi-LLM Systems
Large Language Models
Agent Collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model Selection
Multi-LLM Collaboration
Diversity-aware Selection
Capability-aware Strategy
Large Language Models
🔎 Similar Papers