Distributed Convoluted Rank Regression for Non-Shareable Data under Non-Additive Losses

📅 2026-02-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of modeling non-additive U-statistic losses—such as those arising in convolutional rank regression—under high-dimensional distributed data settings. The authors propose a Distributed Convolutional Rank Regression (DCRR) framework that extends the surrogate likelihood principle to this non-additive context by constructing a tailored surrogate objective. DCRR employs a two-stage sparse estimation procedure combining ℓ₁-penalized iterations with refinement via folded-concave penalties. Theoretically, DCRR achieves the centralized oracle convergence rate with only O(log N) communication rounds, accommodates up to M = o(N/(s² log p)) machines, and enjoys the strong oracle property, non-asymptotic error bounds, and consistent model selection via DHBIC. Empirical results demonstrate its marked superiority over naive divide-and-conquer approaches, particularly under heavy-tailed errors.

Technology Category

Search and Optimization: Non-convex OptimizationReasoning under Uncertainty: Stochastic OptimizationMachine Learning: Distributed Machine Learning & Federated Learning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
We study high-dimensional rank regression when data are distributed across multiple machines and the loss is a non-additive U-statistic, as in convoluted rank regression (CRR). Classical communication-efficient surrogate likelihood (CSL) methods crucially rely on the additivity of the empirical loss and therefore break down for CRR, whose global loss couples all sample pairs across machines. We propose a distributed convoluted rank regression (DCRR) framework that constructs a similar surrogate loss and demonstrate its validity under the non-additive losses. We show that this surrogate shares the same population minimizer as the full-data CRR loss and yields estimators that are statistically equivalent to centralized CRR. Building on this, we develop a two-stage sparse DCRR procedure -- an iterative $\ell_1$-penalized stage followed by a folded-concave refinement -- and establish non-asymptotic error bounds, a distributed strong oracle property, and a DHBIC-type criterion for consistent model selection. A scaling result shows that the number of machines may diverge as $M = o({N/(s^2\log p)})$ while achieving centralized oracle rates with only $O(\log N)$ communication rounds. Simulations and a large-scale real data example demonstrate substantial gains over naive divide-and-conquer, particularly under heavy-tailed errors.
Problem

Research questions and friction points this paper is trying to address.

distributed learning
non-additive loss
rank regression
high-dimensional statistics
U-statistics
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributed learning
non-additive loss
convoluted rank regression
surrogate loss
communication efficiency