🤖 AI Summary
This work addresses the failure of fixed aggregation weights due to client heterogeneity in federated learning and the inability of conventional methods to leverage population-level information for error energy recovery. We propose CROWD, an algorithm grounded in well-conditioned linear inversion that extracts client divergence from optimization trajectories at zero additional cost and dynamically computes debiased aggregation weights via mixed-effects modeling. Theoretically, we establish the first proof that second-order moments can be precisely identified from population divergence, overcoming limitations of random-effects meta-analysis and attaining the Bayesian minimax lower bound. Empirically, evaluated on real-world medical imaging data, CROWD achieves oracle excess risk, significantly outperforming uniform weighting and noise-variance-based strategies.
📝 Abstract
A federated objective is a weighted sum of client risks, and the weights are almost always fixed in advance. We treat them instead as the only instrument of a wisdom-of-crowds mechanism: clients are noisy views of one truth, each seeing it through an independent distortion that is unbiased across the crowd. That the optimal weights are inversely proportional to the clients'error energies is classical; we begin at the question that answer presupposes, which energies belong there and whether a crowd can recover them from itself. Excess risk on the truth is of the exact order of the aggregate bias energy, so no optimizer can repair a bad weight vector; the truth itself is identifiable only up to a linear tilt, so subtracting estimated client biases provably reproduces uniform weighting. The expected per-client second moments, however, are exactly identified from the law of the crowd's disagreement by a well-conditioned linear inversion, a step random-effects meta-analysis cannot take because a source reports once; their realized counterparts are estimable up to an incoherence floor the algorithm can measure. This yields CROWD, which reads the disagreement off the optimization trajectory at no extra cost and matches a Bayesian minimax lower bound in the same constant: per instance as the horizon grows, and unconditionally as the prior becomes diffuse. For arbitrary distortions it stays competitive with the optimal weights, at a ratio governed by a geometric incoherence the algorithm can measure. On real scans split into sites with their own miscalibrated detectors it attains the oracle excess risk; on a companion federation that pulls bias and noise apart, weighting by noise variance is worse than not weighting at all, and CROWD is not.