Score
Designs and implements regularizers, loss terms, priors, and sampling procedures that encourage multiple distinct outputs or diverse latent representations—methods include determinantal point process (DPP) regularization, distance-weighted diversity losses, and diversity-enhanced sampling. Analyzes and tunes the trade-off between output diversity and fidelity or other task objectives during training and inference.
This work addresses the challenge of efficiently selecting diverse, high-quality subsets from massive candidate sets, a task hindered by the NP-hardness and superlinear complexity of traditional DPP-MAP approaches, which struggle to scale beyond millions of items. The authors innovatively reformulate DPP-MAP as a continuous optimization problem on the Stiefel manifold and introduce, for the first time, a nonlinear eigenvalue problem with eigenvector dependency (NEPv). They develop a self-consistent field (SCF) iterative solver that guarantees local convergence under a spectral gap condition. By leveraging low-rank kernel structures and efficient matrix-vector multiplications, the proposed algorithm achieves a time complexity of $O((ndk + nk^2)t)$, enabling near-linear scalability with respect to the candidate set size $n$ and substantially overcoming the scalability limitations of existing combinatorial optimization methods.
Traditional diversity methods in text-to-image retrieval neglect application-specific contextual information and struggle to accommodate multi-attribute requirements. To address this, we propose a novel task—Context-aware Diversity Optimization for Composite Attributes (CDR-CA)—the first to incorporate context awareness into multi-attribute diversity modeling. Methodologically, we formulate a unified manifold-space multi-source Determinantal Point Process (DPP) model and introduce a tangent normalization mechanism to dynamically encode contextual signals, enabling controllable and adaptive diversity regulation. Experiments demonstrate significant improvements in both relevance and practical utility of retrieved results across multiple diversity metrics, effectively balancing diversity and accuracy. The implementation is publicly available.
This work studies weighted least-squares function approximation in $L^2$ space based on random sampling: given an $m$-dimensional subspace $V_m$, how to achieve near-optimal $L^2$ approximation error with minimal sampling cost. We propose a generalized volume resampling framework that, for the first time, achieves expected near-optimal $L^2$ error—i.e., bounded by a constant multiple of the best approximation error—using only $O(m log m)$ samples. Furthermore, in embedding normed spaces, we establish almost-sure error control in the $H$-norm. Our method integrates projection determinantal point processes (DPPs), generalized volume sampling, and independent repeated DPP sampling, significantly enhancing sample diversity and feature selection efficiency. Numerical experiments demonstrate that our approach attains accuracy comparable to i.i.d. or classical volume sampling—but with substantially fewer samples.
This work addresses the shortcut learning problem—where models rely on spurious, easily learnable but unreliable cues due to dataset biases. We propose DiffDiv, a novel framework that leverages diffusion probabilistic models (DPMs), uncovering for the first time their unsupervised feature disentanglement capability during mid-training stages to generate counterfactual samples with novel feature compositions. DiffDiv then integrates diversity regularization with disagreement-based ensemble learning to encourage robust feature acquisition. Crucially, it requires no additional annotations, auxiliary data, or explicit causal assumptions—only unsupervised diversity optimization suffices to mitigate shortcut dependencies. Extensive experiments across multiple benchmarks demonstrate that DiffDiv significantly improves model generalization and robustness; its diversity and performance match or surpass those of state-of-the-art methods relying on auxiliary data.
Existing post-training methods for generative models—such as RLHF and DPO—rely on pairwise preference comparisons over single samples, limiting their ability to model population-level properties like diversity and bias. This work proposes the first preference optimization framework based on *multi-sample* comparisons, introducing two novel algorithms: mDPO and mIPO. These methods directly optimize collective characteristics of generated outputs at the set level, extending DPO and IPO with intra-group consistency constraints and noise-robust mechanisms. Experiments demonstrate that the proposed framework significantly outperforms single-sample baselines in enhancing output diversity, mitigating bias, and maintaining robustness under label noise. The results validate both the effectiveness and necessity of multi-sample comparison for modeling and optimizing population-level behavioral traits in generative models.
Existing subsampling methods based on Determinantal Point Processes (DPPs) struggle to construct continuous DPPs that simultaneously achieve favorable variance reduction properties and lack efficient, structure-preserving discretization schemes. This work proposes a novel wavelet-based continuous DPP and introduces a general discretization framework that converts continuous kernels into low-rank discrete kernels while preserving their variance decay characteristics. The approach is the first to enable DPP-based subsampling for target functions with arbitrarily low regularity and provides explicit convergence rates that depend on the function’s smoothness. The proposed wavelet DPP outperforms existing methods both theoretically and empirically in terms of accuracy and efficiency, substantially broadening the applicability and effectiveness of DPPs in machine learning subsampling tasks.
This work addresses the lack of efficient decoding methods for controlling intra-batch diversity in discrete diffusion models for text generation. The authors propose D5P4, a parallel beam search framework that, for the first time, integrates Determinantal Point Processes (DPPs) into discrete diffusion decoding. By modularizing the beam selection objective, D5P4 explicitly balances generation probability and diversity with negligible computational overhead. The approach combines maximum a posteriori (MAP) inference for DPPs with a scalable greedy solver and supports multi-GPU parallel decoding. Experimental results demonstrate that D5P4 significantly enhances output diversity in both free-form text generation and question-answering tasks while maintaining generation quality on par with strong baselines.
This work addresses a pervasive bias in deep generative models—their systematic underestimation of data diversity after training. For the first time, the study identifies the entropy-based origin of this bias through the lens of finite-sample statistical properties. To quantify the diversity gap between generated samples and real data, the authors employ reference-free diversity metrics, including Vendi and RKE scores. Building on this insight, they propose a diversity-aware regularization and guidance strategy. Extensive experiments across multiple benchmark datasets demonstrate that the proposed approach significantly enhances the diversity of generated samples and effectively mitigates the issue of diversity collapse.
本文提出使用确定性点过程来增强联邦学习中客户端调度的多样性,通过自适应确定性客户端调度(ADCS)方法平衡质量和多样性,以应对异质性并改善最差客户端性能。
This work addresses the limitations of existing image exploration tools, which overly rely on similarity-based ranking and thereby constrain designers’ holistic perception of visual space and pattern discovery during early-stage ideation. The authors propose an interactive exploration prototype that supports gradual adjustment of diversity, introducing for the first time a dynamic interface for explicitly negotiating the trade-off between diversity and similarity—departing from conventional static ranking paradigms. Built upon Determinantal Point Processes (DPPs), the system enables controllable diversity-aware sampling and exposes the underlying tuning mechanism through an intuitive interface. User studies demonstrate that this approach significantly reduces backtracking behavior, facilitates visual discovery, and outperforms existing baseline tools during the initial phases of creative exploration.