🤖 AI Summary
This study aims to decipher the pan-cancer mutational landscape across 43 cancer types to uncover shared molecular mechanisms. Method: We propose the first multi-scale contrastive learning framework for cohort-level cancer clustering, integrating mutation features at both the gene level (COSMIC coding mutations) and chromosome level. The framework employs a TabNet encoder and NT-Xent loss to enable dual-view representation learning, operating entirely in an unsupervised manner—without reliance on labeled data—to construct an interpretable and scalable pan-cancer embedding space. Contribution/Results: Experimental results demonstrate that the learned clusters strongly align with known mutational processes (e.g., APOBEC, homologous recombination deficiency) and tissue-of-origin annotations. The method successfully identifies cross-cancer shared driver patterns and candidate oncogenic pathways, establishing a novel paradigm for pan-cancer molecular subtyping and mechanistic dissection.
📝 Abstract
Motivation. Understanding the pan-cancer mutational landscape offers critical insights into the molecular mechanisms underlying tumorigenesis. While patient-level machine learning techniques have been widely employed to identify tumor subtypes, cohort-level clustering, where entire cancer types are grouped based on shared molecular features, has largely relied on classical statistical methods.
Results. In this study, we introduce a novel unsupervised contrastive learning framework to cluster 43 cancer types based on coding mutation data derived from the COSMIC database. For each cancer type, we construct two complementary mutation signatures: a gene-level profile capturing nucleotide substitution patterns across the most frequently mutated genes, and a chromosome-level profile representing normalized substitution frequencies across chromosomes. These dual views are encoded using TabNet encoders and optimized via a multi-scale contrastive learning objective (NT-Xent loss) to learn unified cancer-type embeddings. We demonstrate that the resulting latent representations yield biologically meaningful clusters of cancer types, aligning with known mutational processes and tissue origins. Our work represents the first application of contrastive learning to cohort-level cancer clustering, offering a scalable and interpretable framework for mutation-driven cancer subtyping.