🤖 AI Summary
This study addresses the high sensitivity of feature selection in untargeted LC-MS metabolomics to preprocessing pipelines, which renders results vulnerable to analytical degrees of freedom. To mitigate this, we introduce multi-universe analysis—a novel, auditable, configuration-driven consensus framework. Building upon a ten-stage quality control filter, our approach integrates four distinct preprocessing strategies with four feature-ranking methods, coupled with bootstrap-based stability selection and label permutation testing to retain only features consistently identified across multiple analytical paths. Among 30,370 initial features, individual pipelines selected 4–20 features with low Jaccard consistency (as low as 0.05), whereas the multi-universe consensus yielded 15 robust features reproducible in at least two out of four pathways. Notably, one feature was stable across all pathways and showed no false positives in 50 permutation tests, substantially enhancing result reliability and reproducibility.
📝 Abstract
Background: Untargeted LC-MS metabolomics requires a long chain of preprocessing decisions, each with several equally defensible options. Analysts typically commit to one pipeline and report the resulting feature shortlist. How strongly that shortlist depends on choices that were never varied stays invisible. Results: We adapt multiverse analysis to untargeted metabolomics feature selection. We present an auditable, configuration-driven pipeline that (i) applies a ten-stage quality-control filter cascade in which every feature's fate is logged, and (ii) runs the downstream analysis as a multiverse over four contrasting preprocessing philosophies, each combined with four feature-ranking methods under bootstrap stability selection and label-permutation testing. Only features recurring across paths enter a tiered consensus. On a demonstration dataset of five breast-cancer cell lines (30,370 detected features), the four single pipelines individually returned shortlists of 4-20 features whose pairwise agreement was as low as Jaccard = 0.05. The multiverse consensus retained 15 features (>=2/4 paths), of which one recurred across all four, although two paths (sharing normalization and drift-correction methods) dominate the consensus. A pipeline-wide label-permutation test found no false discoveries in 50 null permutations. Conclusions: Reporting only preprocessing-robust features, with a complete kept/dropped audit trail, converts hidden analytical degrees of freedom into an explicit, inspectable output. We discuss scope and limitations, including single-batch design and the need for independent validation.