🤖 AI Summary
This work addresses the challenge of effectively integrating the domain-specific expertise of multiple large language models to achieve collective intelligence. The authors propose the Sakana Fugu family of coordinator models, which establish a dynamic, query-adaptive multi-agent collaboration framework that dynamically orchestrates and synthesizes the capabilities of diverse foundation models in real time based on user requests. By combining large-scale fine-tuning, evolutionary algorithms, and reinforcement learning within a dedicated training infrastructure, the project delivers both a production-grade system (Fugu) optimized for low latency and high performance, and a high-accuracy variant (Fugu-Ultra). The approach achieves state-of-the-art results among publicly available models on several challenging benchmarks, including SWE-Bench Pro, Terminal Bench, LiveCodeBench, and GPQA-Diamond.
📝 Abstract
The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. This raises a natural next objective: how to combine the individual specializations of various LLMs into a collectively intelligent system. To this end, we report the development of Sakana Fugu, a family of orchestrator models that harness and amplify the capabilities of an LLM agent team. Fugu models are themselves language models trained to understand user queries and dynamically devise agentic scaffolds to solve them. Through these adaptive scaffolds, Fugu accesses performance beyond any individual LLM agent, achieving state-of-the-art results compared to other publicly accessible models across a range of challenging tasks, including SWE-Bench Pro, Terminal Bench, LiveCodeBench, GPQA-Diamond, Humanity's Last Exam, and CharXiv Reasoning. We release two models: Fugu, which balances performance with latency for everyday use, and Fugu-Ultra, which prioritizes answer quality on the hardest problems. We describe our training paradigm, which encompasses large-scale fine-tuning, evolutionary algorithms, and reinforcement learning approaches, along with the infrastructure and core design principles that turn these methods into a production system. We hope this report encourages further research into multi-agent systems and dynamic, query-adaptive agentic scaffolds as a path toward the next frontier of AI capabilities, accessed through collective intelligence.