Spectral DPPs via NEPv: A Scalable Continuous Relaxation of Determinantal MAP for Diversity-Aware Data Selection

📅 2026-06-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently selecting diverse, high-quality subsets from massive candidate sets, a task hindered by the NP-hardness and superlinear complexity of traditional DPP-MAP approaches, which struggle to scale beyond millions of items. The authors innovatively reformulate DPP-MAP as a continuous optimization problem on the Stiefel manifold and introduce, for the first time, a nonlinear eigenvalue problem with eigenvector dependency (NEPv). They develop a self-consistent field (SCF) iterative solver that guarantees local convergence under a spectral gap condition. By leveraging low-rank kernel structures and efficient matrix-vector multiplications, the proposed algorithm achieves a time complexity of $O((ndk + nk^2)t)$, enabling near-linear scalability with respect to the candidate set size $n$ and substantially overcoming the scalability limitations of existing combinatorial optimization methods.
📝 Abstract
Selecting a small, diverse, high-quality subset from a massive pool of candidates is a recurring primitive in modern machine learning -- data curation and coreset selection for training and fine-tuning large models, active-learning batch acquisition, prompt and exemplar selection for in-context learning, retrieval diversification, and experimental design. Determinantal Point Processes (\DPP s) give a principled, well-calibrated notion of diversity for this task, but their \emph{MAP} objective -- pick a size-$k$ subset $S$ maximizing $\logdet(L_S)$ -- is NP-hard, and the standard greedy and sampling algorithms scale superlinearly in the ground-set size $n$. This cost is prohibitive precisely in the data-centric regime where diversity matters most, where $n$ ranges over millions to billions of candidate examples, features, or embeddings. We recast \DPP-MAP as a continuous optimization problem over the Stiefel manifold, and show that its first-order optimality conditions form a \emph{Nonlinear Eigenvalue Problem with eigenvector dependency} (\NEPv) of a previously unstudied form. This \NEPv\ admits a self-consistent field (\SCF) iteration with a spectral-gap-based local contraction guarantee, giving a principled iterative solver where the diversity objective drives an eigenvector-dependent operator. The resulting algorithm, \OurMethod, requires only matrix-vector products with the kernel and runs in time $O\!\big((ndk+nk^2)\,t\big)$ for a small number of iterations $t$, scaling near-linearly in $n$ and integrating directly with low-rank and feature-map kernels common in ML. This paper focuses on the relaxation, solver, and scaling analysis; full real-data benchmarking is left to a planned empirical study.
Problem

Research questions and friction points this paper is trying to address.

Determinantal Point Processes
DPP-MAP
diversity-aware selection
scalability
NP-hard optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Determinantal Point Processes
Nonlinear Eigenvalue Problem with eigenvector dependency
Stiefel manifold optimization
Self-Consistent Field iteration
Scalable diversity-aware selection
🔎 Similar Papers
2021-06-14IEEE Transactions on Visualization and Computer GraphicsCitations: 12
2024-09-05Citations: 0
💼 Related Jobs
No related jobs found.
R
Richard Yi Da Xu
Hong Kong Baptist University, TadReamk Limited