Token-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM-based multi-agent systems suffer from high token overhead and latency due to the tight coupling of coordination and reasoning, limiting their scalability. This work proposes S1-MAS, a framework that introduces a novel System One-guided computational division mechanism. It decouples coordination from reasoning by delegating bounded coordination decisions to a lightweight controller while offloading complex reasoning to LLMs. Combined with a compact reader and a decision-evidence loop, S1-MAS enables adaptive collaboration without task-specific training. Evaluated across seven benchmarks, the framework achieves superior accuracy while reducing GPT-4o token consumption by 44.9%–97.2% and end-to-end latency by 37.8%–93.0%.
📝 Abstract
Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Systems
Token Efficiency
Scalability
Large Language Models
Coordination Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent Systems
Token Efficiency
System One
Computational Division of Labor
Decision-Evidence Loop
🔎 Similar Papers
No similar papers found.