🤖 AI Summary
This study addresses the high-dimensional heterogeneity between raters and items in ranking data by proposing a Bayesian latent block co-clustering model under the Plackett–Luce observation framework. The model enables items within each rater cluster to share strength parameters, yielding a compact C×K structured representation. Innovatively, it incorporates the Gnedin prior to automatically infer the number of clusters and blocks, and employs a split-merge MCMC strategy to enhance posterior exploration efficiency. In simulation experiments, the method accurately recovers the underlying latent structure. When applied to TCGA pan-cancer gene expression ranking data, it successfully uncovers sample clusters driven by tissue of origin, with biological validity confirmed through gene set enrichment analysis, demonstrating both interpretability and effective data compression.
📝 Abstract
We introduce a Bayesian latent block model that jointly partitions assessors and items under a Plackett--Luce observation model. Assessors are assigned to $C$ clusters and items to $K$ blocks; items in a block share a common strength parameter within each assessor cluster, yielding a parsimonious $C\times K$ co-clustering representation. Independent Gnedin priors infer $C$ and $K$. Data augmentation gives conjugate Gibbs updates and a tractable MCMC sampler with split-merge moves. Simulations characterize recovery and posterior uncertainty as signal, ranking depth, and group balance vary. Applied to the cancer gene atlas (TCGA) pan-cancer top-500 gene-expression rankings, the model reveals tissue-driven sample structure while compressing gene-level heterogeneity into interpretable blocks. Rank-based GSEA of posterior gene scores supports biological interpretation.