🤖 AI Summary
This work addresses the reproducibility and interpretability challenges in existing large language model–based multi-agent systems for discrete-choice tasks, which often suffer from stochastic routing or heuristic coordination strategies. The authors propose a deterministic multi-agent coordination framework following a “multiple analyses, single decision” paradigm: multiple heterogeneous base models independently produce structured analyses, which are then decomposed and aggregated by a dedicated fusion agent using fixed rules to ensure determinism and transparency. To enhance performance without compromising training flexibility, the framework incorporates an exponential moving average (EMA)–guided dynamic agent selection mechanism. Experimental results demonstrate that the method achieves accuracy improvements of over 10 points on MMLU-Pro and more than 50 points on GSM8K, substantially outperforming both single-model and majority-voting baselines, with the EMA-based routing yielding an additional 0.7–2.0 point gain (statistically significant under McNemar’s test).
📝 Abstract
Multi-agent/ensemble approaches can improve discrete-choice reasoning with large language models, but common orchestration methods are often non-deterministic, expensive, and difficult to reproduce. We propose ORCH, a deterministic multi-agent orchestrator that targets higher accuracy and better cost–performance via stable routing.
ORCH uses a pool of heterogeneous LLM agents and a deterministic routing mechanism based on exponential moving average (EMA) performance tracking. For each question, ORCH selects a small subset of agents, obtains candidate answers, and merges them through a controlled aggregation procedure. We evaluate ORCH on multiple discrete-choice benchmarks and compare against single-model baselines and non-routed ensemble strategies under consistent prompting and scoring.
ORCH delivers consistent accuracy improvements over the best low-cost single model and provides additional gains over high-cost single-model baselines on several tasks, while reducing reliance on always-invoking expensive models. The deterministic routing and merge pipeline improves stability across runs.
ORCH demonstrates that deterministic EMA-guided routing can offer a practical and reproducible orchestration strategy for discrete-choice reasoning. This framework can be extended to additional tasks, agent pools, and preference-aware routing policies in future work.