Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses how independently trained neural network populations spontaneously develop effective division of labor without shared gating mechanisms or cross-agent gradients. To this end, it proposes the DISCO mechanism, which integrates DINOv3 feature mask prediction, self-supervised fine-tuning, and distributed collaborative algorithms to enable decentralized local routing and gradient-free communication. This approach allows networks to leverage emergent expertise solely through forward-pass interactions. The contributions demonstrate that specialized division of labor outperforms a single generalist model; notably, random routing achieves 98% accuracy in expert selection, and the collaborative performance remains robust and persistent across diverse conditions.
📝 Abstract
Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using masked prediction of frozen DINOv3 features. We measure specialization by asking whether the best agent for an input aligns with its latent domain, and utilization by asking whether responsibility is distributed across agents. We progressively remove central control, ending with DISCO (DIStributed COllaboration) where each agent locally selects a helper, reads its internal state through a gradient free channel, and rewards its router only for the improvement that help provides. Specialization emerges and is useful. Randomly routed populations underperform a single generalist, while semantically routed populations outperform it, showing that specialization rather than population size drives the gain. Specialization persists without a central router, and gradient free communication lets nonexperts exploit emergent expertise. In DISCO, a random agent helped by the expert matches the solo generalist, while experts surpass it, including on data outside the specialization mixture. Local routers select the emergent expert for 98% of inputs. These effects persist across population size, model capacity, data imbalance, and finetuning seeds, providing measurable evidence for the dynamics needed by decentralized predictive pretraining.
Problem

Research questions and friction points this paper is trying to address.

emergent specialization
self-supervised learning
decentralized collaboration
division of labor
mixture of experts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decentralized Collaboration
Emergent Specialization
Gradient-free Communication
Self-supervised Learning
Local Routing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Aram Davtyan
University of Bern, Computer Vision Group, Institute of Computer Science, Switzerland
P
Pablo Acuaviva
University of Bern, Computer Vision Group, Institute of Computer Science, Switzerland
S
Sebastian Stapf
University of Bern, Computer Vision Group, Institute of Computer Science, Switzerland
Paolo Favaro
Paolo Favaro
Professor of Computer Vision, University of Bern
computer visionmachine learningcomputational photographyinverse problemsoptimization methods