A Multi-Evidence Framework Rescues Low- Power Prognostic Signals and Rejects Statistical Artifacts in Cancer Genomics

📅 2025-10-21
📈 Citations: 0
Influential: 0
📄 PDF

career value

194K/year
🤖 AI Summary
In cancer genomics, low-statistical-power cohorts often yield false-negative driver genes and false-positive passenger genes due to stringent multiple-testing correction. To address this, we propose a multi-evidence integration framework that transcends conventional p-value–driven approaches by jointly leveraging causal inference techniques—specifically inverse probability weighting and doubly robust estimation—and multidimensional biological evidence, including mutational signatures, expression perturbations, and literature-supported functional annotations. We formalize these components into a five-criteria decision model. Our framework markedly improves robustness and specificity in detecting weak oncogenic signals. Applied to TCGA breast cancer data, it successfully identifies *KMT2C* as a candidate driver gene while excluding well-known false positives such as *RYR2*. The method is both interpretable—enabling transparent evidence tracing—and scalable to diverse genomic data types. It establishes a novel paradigm for driver event discovery in small-sample cancer cohorts.

Technology Category

Application Category

📝 Abstract
Motivation: Standard genome-wide association studies in cancer genomics rely on statistical significance with multiple testing correction, but systematically fail in underpowered cohorts. In TCGA breast cancer (n=967, 133 deaths), low event rates (13.8%) create severe power limitations, producing false negatives for known drivers and false positives for large passenger genes. Results: We developed a five-criteria computational framework integrating causal inference (inverse probability weighting, doubly robust estimation) with orthogonal biological validation (expression, mutation patterns, literature evidence). Applied to TCGA-BRCA mortality analysis, standard Cox+FDR detected zero genes at FDR<0.05, confirming complete failure in underpowered settings. Our framework correctly identified RYR2 - a cardiac gene with no cancer function - as a false positive despite nominal significance (p=0.024), while identifying KMT2C as a complex candidate requiring validation despite marginal significance (p=0.047, q=0.954). Power analysis revealed median power of 15.1% across genes, with KMT2C achieving only 29.8% power (HR=1.55), explaining borderline statistical significance despite strong biological evidence. The framework distinguished true signals from artifacts through mutation pattern analysis: RYR2 showed 29.8% silent mutations (passenger signature) with no hotspots, while KMT2C showed 6.7% silent mutations with 31.4% truncating variants (driver signature). This multi-evidence approach provides a template for analyzing underpowered cohorts, prioritizing biological interpretability over purely statistical significance. Availability: All code and analysis pipelines available at github.com/akarlaraytu/causal- inference-for-cancer-genomics
Problem

Research questions and friction points this paper is trying to address.

Rescues low-power prognostic signals in cancer genomics studies
Distinguishes true biological drivers from statistical artifacts
Addresses false negatives and false positives in underpowered cohorts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates causal inference with biological validation
Distinguishes true signals from artifacts via mutation patterns
Prioritizes biological interpretability over statistical significance