CollabFlow: Recursive Self-Improvement of Agent Collaboration

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of error propagation and reward concentration caused by predefined topologies in multi-agent collaboration by proposing a recursive self-improving learning framework. Methodologically, it constructs an architecture comprising a trainable collaborative director and frozen executors, designs an evidence-conditioned communication protocol for dynamic team formation, and introduces a collaborative trajectory balancing objective to ensure policy diversity. Technically, the framework integrates flow matching objectives with graph neural networks to optimize cross-task collaboration. Experimental results demonstrate that the proposed method comprehensively outperforms existing baselines across twelve datasets, with performance improving consistently over successive iterations.
📝 Abstract
Recursive self-improvement (RSI) lets a system improve from its own outcomes; in LLM-based multi-agent systems, Agents refine one another within a task, and outcomes improve how they collaborate across tasks. However, existing multi-agent collaboration leaves this loop open: collaboration is pre-defined at the operator level, topology-only learning keeps verbatim exchange that propagates errors, and reward maximization on a system's own outcomes concentrates on a few teams. To address these challenges, we propose CollabFlow, an RSI system of Learned Agent Collaboration: a trainable Collab-Director constructs teams of complete Agents, a frozen executor runs them, and each round's outcomes retrain the director. Within each round, the edges of a collaboration graph carry protocols of Evidence-Conditioned Communication: a receiver adopts a differing answer only when the sender's evidence is stronger by a margin, so the director learns who communicates and how. Across rounds, we further propose Collaborative Trajectory Balance (CTB), a flow-based objective that credits each team once across its construction orders and targets a reward-proportional distribution over teams, so several good teams stay in play. We also bound how far this self-generated target moves between rounds, which shrinks as records accumulate. On twelve datasets, CollabFlow outperforms all baselines and keeps improving across rounds. Code is available at https://anonymous.4open.science/r/CollabFlow-631E.
Problem

Research questions and friction points this paper is trying to address.

Recursive self-improvement
Multi-agent collaboration
Large language models
Error propagation
Team diversity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improvement
Multi-Agent Collaboration
Evidence-Conditioned Communication
Collaborative Trajectory Balance
Collab-Director