An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear scaling benefits and optimal architecture selection for small-model multi-agent teams. We systematically evaluate diverse orchestration architectures on executable code benchmarks and introduce a novel "generate-transform" decomposition framework that decouples accuracy gains into coverage dividends and transformation efficiency, overcoming the limitations of single aggregate metrics. Our findings reveal that team scaling effects are highly task-dependent. For instance, the Proposer-Critic architecture significantly enhances precision on arithmetic tasks, demonstrating that scaling small-model teams is not a universal lever but rather a targeted strategy jointly constrained by task characteristics and architectural design.
📝 Abstract
Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual task-averaged number conceals. Proposer-Critic captures the arithmetic gains, scaling steepest and, in aggregate, surpassing every other architecture at the largest budget (item-clustered intervals exclude zero), though it ranks among the weakest elsewhere, and no architecture wins across tasks. We explain these trajectories with an exact generate-transform decomposition. Partitioning any workflow into proposal coverage and a downstream transform, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. The decomposition diagnoses each task: arithmetic offers coverage headroom that a critic-guided transform converts, whereas the multiple-choice benchmarks either saturate in coverage or fail to convert it, and on open-ended code generative recovery nearly vanishes so accuracy tracks coverage. At equal call budgets token cost still varies 2.1x. Extra calls therefore create candidate opportunity that only some architectures, on some tasks, convert. Team scaling is a task- and architecture-specific bet, not a uniform lever.
Problem

Research questions and friction points this paper is trying to address.

multi-agent scaling
orchestration architectures
small LLMs
task-dependent returns
team collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generate-Transform Decomposition
Agent Orchestration
Team Scaling
Proposal Coverage
Proposer-Critic
🔎 Similar Papers
2024-08-10AAAI Conference on Artificial IntelligenceCitations: 30