determinantal point process sampling

Designs and implements algorithms and procedures that select diverse, non‑redundant subsets from a larger set by modeling repulsion with a determinantal point process; this includes constructing kernels, batch‑sampling routines, and mechanisms to trade off element quality and diversity. Builds or analyzes methods for efficient DPP sampling, batch acquisition, and evaluation of coverage/representation to reduce redundancy among selected samples.

determinantalpointprocesssampling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of efficiently selecting diverse, high-quality subsets from massive candidate sets, a task hindered by the NP-hardness and superlinear complexity of traditional DPP-MAP approaches, which struggle to scale beyond millions of items. The authors innovatively reformulate DPP-MAP as a continuous optimization problem on the Stiefel manifold and introduce, for the first time, a nonlinear eigenvalue problem with eigenvector dependency (NEPv). They develop a self-consistent field (SCF) iterative solver that guarantees local convergence under a spectral gap condition. By leveraging low-rank kernel structures and efficient matrix-vector multiplications, the proposed algorithm achieves a time complexity of $O((ndk + nk^2)t)$, enabling near-linear scalability with respect to the candidate set size $n$ and substantially overcoming the scalability limitations of existing combinatorial optimization methods.

Determinantal Point Processesdiversity-aware selectionDPP-MAP

Existing subsampling methods based on Determinantal Point Processes (DPPs) struggle to construct continuous DPPs that simultaneously achieve favorable variance reduction properties and lack efficient, structure-preserving discretization schemes. This work proposes a novel wavelet-based continuous DPP and introduces a general discretization framework that converts continuous kernels into low-rank discrete kernels while preserving their variance decay characteristics. The approach is the first to enable DPP-based subsampling for target functions with arbitrarily low regularity and provides explicit convergence rates that depend on the function’s smoothness. The proposed wavelet DPP outperforms existing methods both theoretically and empirically in terms of accuracy and efficiency, substantially broadening the applicability and effectiveness of DPPs in machine learning subsampling tasks.

CoresetDeterminantal Point ProcessesMinibatch

Candidate set sampling: A note on theoretical guarantees

Dec 12, 2025
SX
Shifeng Xiong
🏛️ Academy of Mathematics and Systems Science | Chinese Academy of Sciences

To address efficient sampling in high-dimensional spaces when only an unnormalized density function is available, this paper proposes Candidate Set Sampling (CSS), a non-iterative, dimension-agnostic, and hyperparameter-free numerical sampling method. CSS directly discretizes the density function to construct a finite candidate set, bypassing conventional frameworks such as Markov Chain Monte Carlo (MCMC) or variational inference. We provide a theoretical guarantee that the induced sampling distribution converges exponentially fast to the target distribution in total variation distance—the first rigorous convergence result for such discretization-based sampling methods. Empirical evaluations demonstrate that CSS maintains low computational overhead and stable accuracy in high dimensions, exhibits rapid convergence, and is straightforward to implement. It significantly outperforms existing black-box sampling methods across diverse benchmarks.

Based on discretization of density functionIntroduces candidate set sampling methodNon-iterative, dimension-free, fast convergence

This study addresses the NP-hard problem of selecting low-discrepancy subsets from large-scale sets, which has significant applications in quasi-Monte Carlo methods, machine learning, and computer graphics. We establish for the first time the NP-hardness of this problem under kernel discrepancy measures and propose a novel framework based on Bayesian optimization. By constructing a surrogate model using deep embedded kernels, our approach efficiently searches for optimal subsets, overcoming the computational bottlenecks inherent in traditional combinatorial optimization. Extensive experiments demonstrate that the method substantially reduces subset discrepancy across multiple discrepancy metrics, highlighting its effectiveness and versatility in low-discrepancy design tasks.

kernel discrepancylow-discrepancyNP-hard

Weighted least-squares approximation with determinantal point processes and generalized volume sampling

Dec 21, 2023
AN
Anthony Nouy
🏛️ Nantes Université | Centrale Nantes | CNRS UMR 6629

This work studies weighted least-squares function approximation in $L^2$ space based on random sampling: given an $m$-dimensional subspace $V_m$, how to achieve near-optimal $L^2$ approximation error with minimal sampling cost. We propose a generalized volume resampling framework that, for the first time, achieves expected near-optimal $L^2$ error—i.e., bounded by a constant multiple of the best approximation error—using only $O(m log m)$ samples. Furthermore, in embedding normed spaces, we establish almost-sure error control in the $H$-norm. Our method integrates projection determinantal point processes (DPPs), generalized volume sampling, and independent repeated DPP sampling, significantly enhancing sample diversity and feature selection efficiency. Numerical experiments demonstrate that our approach attains accuracy comparable to i.i.d. or classical volume sampling—but with substantially fewer samples.

Approximating L2 functions using m-dimensional space V_mPromoting feature diversity via DPP and volume samplingReducing sample count while maintaining error bounds

Latest Papers

What's happening recently
View more

该研究针对选区重划中的高维组合问题,通过构建和诊断具有适当覆盖和重叠的候选层来实现分层抽样,使用聚类方法形成计划层面的'单词'。

balanced graph partitionscandidate strataphase space

This study addresses the problem of clustering probability distributions with repulsion to uncover semantically clear and well-separated data structures. To this end, the authors propose the distributed Determinantal Point Process (dDPP), which extends Determinantal Point Processes to the space of probability distributions for the first time. Treating distributions as atomic elements, they construct an L-ensemble using a sliced Wasserstein kernel and embed it within a generalized Bayesian mixture model to induce random partitions. The method innovatively incorporates a repulsive mechanism among distributions and introduces a utility function based on hierarchical optimal transport for posterior summarization. Experiments on single-cell gene expression and human epilepsy datasets demonstrate that the approach effectively reveals intrinsic structures, yielding highly separable and interpretable clustering results.

determinantal point processdistribution-valued clusteringprobability distributions

This work addresses the challenge of effectively balancing sample representativeness and diversity in dynamic data selection to accelerate training while preserving model accuracy. The authors propose a novel framework that defines representativeness as coverage of high-frequency feature factors in the dataset, while diversity is achieved by progressively introducing complementary rare factors during training. Leveraging sparse autoencoders, the method identifies sparsely activated units in the feature space and estimates both sample-level and dataset-level factor distributions. A frequency-based penalty mechanism combined with a smooth scheduling strategy enables efficient, gradient-free data selection. Evaluated across five vision and language benchmarks, the approach matches or exceeds the performance of full-data training while achieving over 2× speedup in training time.

diversitydynamic data selectionrepresentativeness

This work addresses geometric data pruning methods that rely on neighborhood similarity assumptions, which inherently introduce selection bias. Discarding this assumption, we reformulate unbiased subset selection from first principles as a variance minimization problem. Through a linear programming perspective, we construct high-dimensional polytopes and derive closed-form pairwise variance expressions, enabling an efficient vertex-walking algorithm for label-agnostic data pruning with strictly guaranteed statistical unbiasedness. Experiments across multiple benchmarks demonstrate that the proposed method outperforms uniform sampling and mainstream geometric approaches in accuracy, exhibiting particularly superior performance under small selection budgets while effectively reducing stochastic gradient descent (SGD) variance.

Dataset PruningLabel-FreeSubset Selection

This work addresses the challenge of maintaining stability in large language model fine-tuning under unreliable conditions—such as storage failures and communication errors—where conventional data selection methods often falter. The authors propose ProbDPP, the first framework to integrate data access reliability into Determinantal Point Processes (DPPs). By introducing a regularized objective that jointly optimizes geometric diversity and reliability-aware costs, the approach formulates data selection as a combinatorial semi-bandit problem and devises a UCB-style online learning algorithm. ProbDPP enables robust selection of diverse data subsets even when reliability information is unknown. Theoretical analysis establishes a bounded regret guarantee, demonstrating significant improvements in both the robustness and deployment efficiency of data selection for fine-tuning.

Data UncertaintyDeterminantal Point ProcessesInformative Data Selection

Hot Scholars

RB

Rémi Bardenet

CNRS, CRIStAL, Ecole Centrale Lille, Univ. Lille, France
Computational statisticsmachine learningapplications to biology and physics
MV

Michal Valko

Chief Models Officer @ Stealth Startup, Inria & MVA - Ex: Llama at Meta; Gemini and BYOL @ Deepmind
large language modelsreasoningfine-tuningtest-time computation
PG

Pranav Gupta

Assistant Professor, Gies College of Business, UIUC
Collective IntelligenceHuman-AI TeamingTransactive AttentionDigital Nudging
GL

Gawon Lee

Pusan National University
sequence modelingsingal processingdeep learning