Score
Designs and analyzes conditioning matrices or data-driven preconditioners built from public (non-sensitive) features to reshape feature spectra and accelerate convergence of private optimization and estimation algorithms. Works to construct these matrices so they provably speed convergence while avoiding additional consumption of privacy budget.
This work addresses the challenge of effectively leveraging public non-sensitive features to improve performance in regression tasks under label differential privacy or with partially sensitive features, where existing methods fall short. The authors propose Cond-DP, a novel approach that, for the first time, incorporates the spectral decay properties of public features into differentially private optimization. By constructing a data-driven conditioning matrix that reshapes the optimization landscape, Cond-DP yields a provably faster-converging variant of differentially private stochastic gradient descent (DPSGD). Crucially, this enhancement incurs no additional privacy cost. Empirical evaluations across diverse datasets and model architectures demonstrate that Cond-DP consistently and significantly outperforms current baselines, exhibiting superior practicality and robustness.
To address the high computational overhead of matrix decomposition-based methods in differentially private model training—particularly the need for numerical solving of pre-optimization problems—this paper proposes Banded Square Root (BSR) matrix decomposition. BSR is the first method to exploit the structural properties of the standard matrix square root to construct a banded approximation, enabling zero-overhead initialization without any pre-optimization. It further yields closed-form analytical solutions for momentum SGD and weight decay. Theoretically, BSR provides guaranteed approximation accuracy in both centralized and federated learning settings under differential privacy. Empirical results demonstrate that BSR achieves model accuracy comparable to optimal decompositions while completely eliminating the traditional numerical optimization step, thereby significantly improving training efficiency—especially at scale.
This work addresses efficiency bottlenecks in solving large-scale linear systems and approximating matrix norms. We propose a multilevel randomized sketching preconditioned iterative method, integrating Nyström low-rank approximation, sparse random sketching, and multilevel preconditioning. It establishes the first multilevel sketched preconditioning framework grounded in the natural average condition number. Theoretical contributions include: (1) optimal complexity $ ilde{O}(n^2 + d_lambda^omega)$ for solving regularized linear systems; (2) accelerated complexity $ ilde{O}(n^{2.065} + k^omega)$ for systems with $k$ outlying singular values; and (3) Schatten-$p$ norm approximation—particularly the nuclear norm—at $ ilde{O}(n^{2.11})$, improving upon the prior best $ ilde{O}(n^{2.18})$. These advances significantly enhance computational efficiency for key subproblems in applications such as Gaussian process regression.
Existing adaptive optimizers (e.g., Adam) neglect the gradient covariance structure, resulting in poor directional adaptivity and slow convergence. This work proposes an implicit diagonalization method based on invertible linear coordinate transformations, which maps the preconditioning matrix into an approximately diagonal space without explicit low-rank or sparsity approximations—thereby balancing computational efficiency and directional modeling fidelity. The approach integrates seamlessly into memory-efficient optimizers (e.g., Adafactor) while preserving full compatibility with standard training pipelines. Evaluated on large language models including LLaMA, it achieves a 2× speedup in convergence and significantly accelerates training across diverse deep models, without increasing GPU memory overhead. The core contribution is the first use of invertible transformations to enable implicit diagonalization of preconditioning matrices, breaking the traditional trade-off between computational cost and accuracy in covariance modeling.
In overparameterized nonconvex matrix factorization—where the specified rank $r$ exceeds the true rank $r^*$—gradient descent suffers sublinear convergence, severely limiting efficiency. This paper proposes PrecGD, a lightweight preconditioned gradient descent method that restores linear convergence without requiring prior knowledge of $r^*$. Key contributions include: (i) the first theoretical demonstration that $ell_2$ regularization, within a specific damping range, effectively mitigates ill-conditioning of the factor matrices; and (ii) a novel adaptive damping strategy, computed cheaply from current iterates, which robustly handles the conditioning of the ground-truth solution. PrecGD maintains linear convergence even under noise and achieves the information-theoretically optimal estimation error bound. Experiments across diverse overparameterized matrix sensing and factorization tasks confirm substantial improvements in both convergence speed and reconstruction accuracy.
This work addresses the geometric mismatch between the isotropic noise injected by DP-SGD and the anisotropic loss landscape of deep neural networks in differentially private optimization. Existing preconditioning methods either consume privacy budget by using private data or suffer from distributional shift when relying on public data. To overcome this, the authors propose a KFAC-based preconditioner that requires no real data: it employs structured synthetic noise to probe the network, decoupling the Fisher information matrix into an architecture-sensitive component—recovered via synthetic noise—and an input-dependent component—approximated using modality-specific spectral statistics. This approach is the first to estimate curvature information without accessing either private or public data. Under strong privacy constraints (ε ≤ 3), DP-KFC consistently outperforms DP-SGD and adaptive baselines across multimodal tasks, matching the performance of private-data-based methods while avoiding the up to 4.8% accuracy drop caused by reliance on public data.
This work addresses the challenge in differentially private federated learning where stringent privacy budgets necessitate large injected noise, severely slowing the convergence of first-order methods, while existing second-order approaches are hindered by prohibitive memory costs in high-dimensional models. To overcome this, the authors propose a server-side second-order optimization framework that constructs a natural gradient preconditioner using the Fisher information matrix and leverages the Sherman-Morrison formula for efficient matrix inversion, requiring only O(d) memory and computational complexity per client. This approach is the first to enable scalable second-order optimization under (ε,δ)-differential privacy, effectively balancing privacy guarantees with convergence efficiency. Experiments on CIFAR-10 demonstrate that the method consistently achieves significantly higher test accuracy than first-order baselines across various privacy budgets.
This work addresses the computational inefficiency of existing Gaussian sketching methods in differentially private linear regression, despite their strong utility guarantees. The authors propose a novel differentially private sketching mechanism based on fast structured transforms—such as the Hadamard transform—to enable efficient data compression in ordinary least squares estimation. Their approach is the first to simultaneously achieve near-optimal accuracy comparable to Gaussian sketches and computational speed approaching that of non-private fast sketching algorithms, all under rigorous differential privacy guarantees. Under typical parameter settings, the method significantly outperforms current differentially private alternatives in runtime while delivering state-of-the-art trade-offs between efficiency and statistical utility.
This work addresses the problem of efficiently approximating the leading principal components of high-dimensional data matrices under $(\varepsilon, \delta)$-differential privacy, where neighboring datasets differ by a single row. We propose a noise-filtering mechanism that adapts to the matrix coherence and is integrated into the power iteration algorithm, significantly improving the accuracy of principal component estimation while preserving privacy. By moving beyond worst-case analyses, our approach achieves particularly strong performance on low-coherence matrices. Furthermore, this method extends the framework of Hardt and Roth from the more restrictive user-level privacy model to the more practical and widely applicable row-level privacy setting.
This work addresses the challenge of balancing memory efficiency and utility in multi-round differentially private training by proposing the γ-BIFR factorization method. Leveraging banded inverse factorization, γ-BIFR constructs an explicit covariance matrix decomposition with a tunable parameter γ, unifying and generalizing existing low-memory, high-bandwidth correlated noise mechanisms. The approach enables flexible noise buffer configurations and significantly outperforms DP-λCGD and BISR under low-bandwidth and low-memory constraints. It achieves higher model utility while reducing both RMSE and amplified RMSE, and provides a tighter theoretical bound on the multi-round participation error.