Score
Designs and implements methods to summarize and postprocess posterior distributions defined on permutations or other discrete ranking spaces, producing consensus-ranking point estimates and marginal posterior summaries (e.g., item inclusion or position probabilities). Builds algorithms that aggregate rank posteriors into representative rankings and quantify uncertainty for those discrete summaries through credible sets, posterior probabilities, or loss-minimizing point estimates.
Bayesian clustering quantifies posterior uncertainty but lacks a general method to summarize uncertainty over the partition space. This paper proposes a generic post-processing framework for Markov chain Monte Carlo (MCMC) posterior samples: it constructs the first interpretable and computationally efficient posterior credible set of clusterings; introduces a novel uncertainty measure enabling cumulative posterior estimation and credible region construction for cluster-specific parameters; and operates entirely without point estimates of partitions or label alignment. By integrating statistical modeling on the partition space with optimization-based aggregation, the method enhances interpretability, robustness, and practical utility across multiple empirical studies. It establishes a reliable, off-the-shelf paradigm for uncertainty quantification in Bayesian clustering.
Bayesian density and clustering modeling inherently face a tension between interpretability and flexibility: overparameterized models achieve high fit accuracy but lack transparency. This paper proposes a posterior-projection summarization framework that, grounded in decision theory, optimally compresses the high-dimensional posterior predictive distribution into a low-dimensional parametric density and clustering representation—preserving the original model’s fit fidelity while quantifying uncertainty. Our approach is the first to systematically integrate nonparametric modeling, posterior dimensionality reduction, and Bayesian uncertainty propagation, thereby jointly optimizing statistical interpretability and fitting fidelity. Experiments on synthetic and real-world datasets demonstrate that the resulting summaries enjoy theoretical guarantees and practical utility: fitting loss remains tightly controlled, and model transparency—as well as downstream interpretability for analysis—is substantially enhanced.
This work addresses the challenge of effectively modeling ranking distributions over the symmetric group by proposing the Consensus Ranking Distribution (CRD) model, which approximates the target distribution via a sparse mixture of Dirac measures. The approach introduces local ranking medians and employs the Kendall τ distance as the optimal transport cost, leveraging a top-down tree-structured algorithm to iteratively refine approximation accuracy. Theoretical analysis demonstrates that the transport distortion can be precisely expressed in terms of pairwise ranking probabilities, thereby circumventing the fundamental obstacle posed by the absence of a vector space structure on the symmetric group. Experimental results validate the superior efficiency and practical utility of the proposed algorithm.
This work addresses the limitation of classical deterministic ranking methods—such as Borda count and Copeland—which ignore uncertainty arising from sampling noise or missing data, often leading to erroneous conclusions. The authors model the true ranking as a latent random variable and introduce a probabilistic ranking framework grounded in pairwise win probabilities, accompanied by an efficient approximate inference algorithm. Their key contribution lies in formally integrating uncertainty quantification into classical ranking paradigms for the first time, proposing the Worst-Best Rank method to construct confidence intervals at both item-wise and global ranking levels. This approach effectively corrects bias induced by missing data, enabling robust estimation of true item performance even under substantial incompleteness, thereby significantly enhancing the reliability, transparency, and fairness of rankings in high-stakes decision-making contexts.
To address the challenge of dimensionality reduction for high-dimensional data in Approximate Bayesian Computation (ABC), this paper proposes a unified, theoretically grounded, and computationally tractable framework for selecting summary statistics—based on minimizing the Expected Posterior Entropy (EPE) under the prior predictive distribution. This criterion is rigorously shown to subsume leading approaches—including SS-ABC, LDA-ABC, and Bayesian Synthetic Likelihood—as special cases or asymptotic limits. The method integrates information-theoretic optimization, prior predictive modeling, dimensionality reduction, and Monte Carlo estimation to enable end-to-end learning of low-dimensional, high-fidelity summaries. Evaluated on benchmark and real-world applications, the proposed strategy consistently improves posterior accuracy and stability, offering a plug-and-play, general-purpose solution for ABC-based inference.
This work addresses the non-uniform distribution of posterior predictive p-values (ppp) under the Bayesian framework, which hinders reliable model diagnostics and cross-model comparisons. The authors propose a natural calibration method that transforms ppp values into calibrated posterior predictive p-values (cppp), which follow a standard uniform distribution under the true model. This calibration establishes, for the first time, a unified and comparable scale for ppp-based assessments. The approach is grounded in a double-simulation computational framework that seamlessly integrates Bayesian inference with posterior predictive checking, and it is applicable to both parametric and nonparametric models. Theoretical analysis demonstrates favorable statistical properties of cppp, while empirical studies illustrate its effectiveness in enabling fair comparisons among models and prior specifications on real-world data.
研究针对带平局的排名提出了一种后处理方法,通过精确算法和快速启发式方法来确定最接近的公平共识排名。
本文提出了一种高效的贝叶斯估计方法,通过引入辅助变量线性化Benter模型中的不可解标准化项,解决了该模型拟合的挑战。
本文提出了一种非参数贝叶斯推理框架,用于解决部分识别的离散响应模型问题,通过直接对条件概率质量函数进行推理,避免了将条件矩转换为无条件矩或离散化协变量的需求。
This study addresses the problem of efficient sequential data collection in political polling, aiming to terminate sampling as early as possible while preserving statistical validity. To this end, the authors propose a posterior optimal sequential testing procedure based on e-values, establishing—for the first time—a statistical testing framework for the Schulze voting system and devising tailored tests for Condorcet and Borda rules. The key innovation lies in solving the challenging construction of optimal e-values under non-convex composite hypotheses with multivariate Bernoulli data, achieved via an efficient Frank–Wolfe algorithm for computing reverse information projection, applicable to general composite null and alternative hypotheses. Empirical evaluations on both synthetic simulations and real-world data from the 2022 French presidential election demonstrate that the proposed method substantially outperforms state-of-the-art approaches in both statistical power and sample efficiency.