posterior summarization for rankings

Designs and implements methods to summarize and postprocess posterior distributions defined on permutations or other discrete ranking spaces, producing consensus-ranking point estimates and marginal posterior summaries (e.g., item inclusion or position probabilities). Builds algorithms that aggregate rank posteriors into representative rankings and quantify uncertainty for those discrete summaries through credible sets, posterior probabilities, or loss-minimizing point estimates.

posteriorsummarizationforrankings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Uncertainty Quantification in Bayesian Clustering

Nov 19, 2025
GL
Garritt L. Page
🏛️ Brigham Young University | Florida State University | Duke University

Bayesian clustering quantifies posterior uncertainty but lacks a general method to summarize uncertainty over the partition space. This paper proposes a generic post-processing framework for Markov chain Monte Carlo (MCMC) posterior samples: it constructs the first interpretable and computationally efficient posterior credible set of clusterings; introduces a novel uncertainty measure enabling cumulative posterior estimation and credible region construction for cluster-specific parameters; and operates entirely without point estimates of partitions or label alignment. By integrating statistical modeling on the partition space with optimization-based aggregation, the method enhances interpretability, robustness, and practical utility across multiple empirical studies. It establishes a reliable, off-the-shelf paradigm for uncertainty quantification in Bayesian clustering.

Generating credible sets for clustering without partition estimatesProviding new uncertainty measures and cluster-specific parameter estimatesSummarizing posterior uncertainty in Bayesian clustering models

Lower-dimensional posterior density and cluster summaries for overparameterized Bayesian models

Jun 11, 2025
HB
Henrique Bolfarine
🏛️ The University of Texas at Austin | Insper

Bayesian density and clustering modeling inherently face a tension between interpretability and flexibility: overparameterized models achieve high fit accuracy but lack transparency. This paper proposes a posterior-projection summarization framework that, grounded in decision theory, optimally compresses the high-dimensional posterior predictive distribution into a low-dimensional parametric density and clustering representation—preserving the original model’s fit fidelity while quantifying uncertainty. Our approach is the first to systematically integrate nonparametric modeling, posterior dimensionality reduction, and Bayesian uncertainty propagation, thereby jointly optimizing statistical interpretability and fitting fidelity. Experiments on synthetic and real-world datasets demonstrate that the resulting summaries enjoy theoretical guarantees and practical utility: fitting loss remains tightly controlled, and model transparency—as well as downstream interpretability for analysis—is substantially enhanced.

Balancing interpretability and flexibility in Bayesian modelsProjecting complex models to lower-dimensional summariesProviding uncertainty quantification for density and cluster estimates

This work addresses the challenge of effectively modeling ranking distributions over the symmetric group by proposing the Consensus Ranking Distribution (CRD) model, which approximates the target distribution via a sparse mixture of Dirac measures. The approach introduces local ranking medians and employs the Kendall τ distance as the optimal transport cost, leveraging a top-down tree-structured algorithm to iteratively refine approximation accuracy. Theoretical analysis demonstrates that the transport distortion can be precisely expressed in terms of pairwise ranking probabilities, thereby circumventing the fundamental obstacle posed by the absence of a vector space structure on the symmetric group. Experimental results validate the superior efficiency and practical utility of the proposed algorithm.

consensus rankingKemeny medianKendall tau distance

This work addresses the limitation of classical deterministic ranking methods—such as Borda count and Copeland—which ignore uncertainty arising from sampling noise or missing data, often leading to erroneous conclusions. The authors model the true ranking as a latent random variable and introduce a probabilistic ranking framework grounded in pairwise win probabilities, accompanied by an efficient approximate inference algorithm. Their key contribution lies in formally integrating uncertainty quantification into classical ranking paradigms for the first time, proposing the Worst-Best Rank method to construct confidence intervals at both item-wise and global ranking levels. This approach effectively corrects bias induced by missing data, enabling robust estimation of true item performance even under substantial incompleteness, thereby significantly enhancing the reliability, transparency, and fairness of rankings in high-stakes decision-making contexts.

deterministic rankingfairness in rankingincomplete data

Unifying Summary Statistic Selection for Approximate Bayesian Computation

Jun 06, 2022
TH
Till Hoffmann
🏛️ Harvard T.H. Chan School of Public Health

To address the challenge of dimensionality reduction for high-dimensional data in Approximate Bayesian Computation (ABC), this paper proposes a unified, theoretically grounded, and computationally tractable framework for selecting summary statistics—based on minimizing the Expected Posterior Entropy (EPE) under the prior predictive distribution. This criterion is rigorously shown to subsume leading approaches—including SS-ABC, LDA-ABC, and Bayesian Synthetic Likelihood—as special cases or asymptotic limits. The method integrates information-theoretic optimization, prior predictive modeling, dimensionality reduction, and Monte Carlo estimation to enable end-to-end learning of low-dimensional, high-fidelity summaries. Evaluated on benchmark and real-world applications, the proposed strategy consistently improves posterior accuracy and stability, offering a plug-and-play, general-purpose solution for ABC-based inference.

Characterizing summary classes for analyzing dimensionality reduction algorithmsSelecting low-dimensional summary statistics for efficient likelihood-free inferenceUnifying framework for obtaining informative summaries via expected posterior entropy

Latest Papers

What's happening recently
View more

This work addresses the non-uniform distribution of posterior predictive p-values (ppp) under the Bayesian framework, which hinders reliable model diagnostics and cross-model comparisons. The authors propose a natural calibration method that transforms ppp values into calibrated posterior predictive p-values (cppp), which follow a standard uniform distribution under the true model. This calibration establishes, for the first time, a unified and comparable scale for ppp-based assessments. The approach is grounded in a double-simulation computational framework that seamlessly integrates Bayesian inference with posterior predictive checking, and it is applicable to both parametric and nonparametric models. Theoretical analysis demonstrates favorable statistical properties of cppp, while empirical studies illustrate its effectiveness in enabling fair comparisons among models and prior specifications on real-world data.

Bayesian inferencecalibrationmodel comparison

This study addresses the problem of efficient sequential data collection in political polling, aiming to terminate sampling as early as possible while preserving statistical validity. To this end, the authors propose a posterior optimal sequential testing procedure based on e-values, establishing—for the first time—a statistical testing framework for the Schulze voting system and devising tailored tests for Condorcet and Borda rules. The key innovation lies in solving the challenging construction of optimal e-values under non-convex composite hypotheses with multivariate Bernoulli data, achieved via an efficient Frank–Wolfe algorithm for computing reverse information projection, applicable to general composite null and alternative hypotheses. Empirical evaluations on both synthetic simulations and real-world data from the 2022 French presidential election demonstrate that the proposed method substantially outperforms state-of-the-art approaches in both statistical power and sample efficiency.

composite hypothesese-valuesnon-convex parameter sets

Hot Scholars

VV

Valeria Vitelli

Associate Professor, Department of Biostatistics, University of Oslo, Norway
preference learningfunctional data analysisclusteringhigh-dimensional data
NF

Nial Friel

Professor of Statistics, University College Dublin.
Bayesian statisticsMonte Carlo methodsstatistical network analysis
PG

Paul Gustafson

Department of Statistics, University of British Columbia
Bayesian methodscausal inferenceevidence synthesismeasurement error
MN

Maria Nareklishvili

Researcher in Econometrics, Stanford University
Econometrics and Data Science
SL

Stephen Lindsay

Lecturer, Glasgow University
Human Computer Interaction