Score
Designs, builds, and evaluates machine‑learning models and systems that produce structured outputs (for example graphs, coordinate sets, or multi‑label/sequence predictions), including structured‑prediction architectures and specialized structure‑prediction models. Work includes implementing uncertainty calibration methods such as conformal prediction, creating models that map structures to functions, building prediction‑serving pipelines, and analyzing model performance, robustness, and failure modes.
This work addresses the challenge of quantifying prediction uncertainty in generative biomolecular design, where feedback covariate shift undermines conventional uncertainty estimation. We propose the first conformal prediction framework tailored to closed-loop design paradigms. Departing from standard i.i.d. assumptions, our method imposes no structural constraints on either the design algorithm or the regression model, delivering finite-sample statistically valid confidence sets for arbitrary black-box design pipelines. Key innovations include quantile-regression-driven adaptive conformal prediction, explicit modeling of feedback-induced distributional shift, and robust error calibration. Evaluated on protein and small-molecule design tasks, our approach achieves ≥94.8% empirical coverage at the 95% nominal confidence level—substantially outperforming standard conformal methods (which drop to as low as 72%)—while maintaining high predictive accuracy.
This paper addresses the challenge of uncontrolled decision risk arising from insufficient identification and quantification of uncertainty in machine learning models. It systematically distinguishes epistemic uncertainty (model ignorance) from aleatoric uncertainty (inherent data noise) and proposes a unified uncertainty modeling framework grounded in conformal prediction. Methodologically, conformal prediction is integrated into linear regression, random forests, and neural networks to produce prediction intervals with rigorous finite-sample statistical guarantees. A key contribution lies in the decoupled, model-agnostic adaptation of conformal prediction to diverse mainstream models—preserving theoretical validity while ensuring practical deployability. Empirical evaluations demonstrate substantial improvements in predictive reliability and interpretability. The framework supports risk-aware decision-making in domains such as financial risk management and operational optimization, providing a verifiable and reproducible technical pathway for uncertainty-driven business decisions.
Existing studies lack rigorous theoretical characterization of posterior calibration methods—such as Platt scaling and isotonic regression—particularly regarding their dependence on feature quality, generalizability across models and datasets, and convergence behavior and robustness under finite-sample regimes. Method: We establish a unified theoretical framework for these two dominant calibration paradigms, deriving the first non-asymptotic guarantees on convergence rates, computational complexity, and explicit sample-size dependencies. Our analysis quantifies the relationship between feature informativeness and calibration robustness. Results: Through synthetic experiments and extensive empirical evaluation across diverse model architectures and benchmark datasets, we validate our theoretical findings. The results yield actionable guidance: isotonic regression is preferable under low signal-to-noise ratios or limited samples, whereas Platt scaling exhibits superior robustness in high-dimensional sparse feature settings. Our work provides interpretable, reusable principles for uncertainty calibration in practical machine learning systems.
In multi-model prediction, system reliability is jointly affected by model uncertainty—arising from statistical dependencies among models trained on shared data—and input uncertainty—stemming from the inherent randomness of inputs. Existing methods fail to simultaneously and rigorously quantify these two distinct uncertainty sources. This paper proposes the first theoretical framework that explicitly decouples inter-model dependencies from input stochasticity, treating them as independent variables and constructing their joint probability distribution. Through rigorous probabilistic modeling and statistical analysis, we derive an analytical characterization of the joint distribution of multi-model outputs. This enables, for the first time, systematic and unified quantification of both model and input uncertainties at the system level. The framework establishes a principled foundation for high-reliability decision-making, robust system design, and subsequent uncertainty propagation algorithms.
In machine learning–guided design, a critical challenge is reliably selecting a design algorithm that satisfies user-specified success criteria—e.g., ensuring ≥10% of generated designs exceed a threshold on a target property. Method: We propose the first prediction-powered inference framework for algorithm selection. It integrates predictions from a learned model with held-out labeled data, employing density-ratio weighting for calibration. The method provides theoretical guarantees that the selected algorithm satisfies the desired success probability constraint with high confidence, supporting both known and estimable density ratios; it either returns the optimal algorithm or certifies infeasibility. Crucially, it jointly optimizes the predictive and generative models. Results: Evaluated on simulated protein and RNA design tasks, our approach significantly improves accuracy in identifying successful algorithms while strictly enforcing user-specified probabilistic constraints—empirically validating the practicality of its theoretical guarantees.
Traditional conformal prediction (CP) provides only marginal coverage guarantees under small-sample calibration, exhibiting high variance in coverage distribution and frequent violations below the nominal level—thereby undermining reliability in uncertainty quantification. To address this, we propose a novel conformal prediction framework that, for the first time, delivers probabilistic coverage guarantees for individual predictors—e.g., $ mathbb{P}( ext{Coverage} geq 1-alpha) geq 1-delta $—overcoming the fundamental limitation of marginal guarantees. This guarantee holds rigorously even with limited calibration data and asymptotically recovers classical CP guarantees under large samples. Our method leverages nonparametric concentration inequalities, requires no assumptions on error distributions, and integrates seamlessly with mainstream CP libraries. Experiments demonstrate substantial improvements in coverage stability and safety under low-data regimes, providing verifiable statistical guarantees for uncertainty quantification in resource-constrained settings.
This work addresses the inefficiencies in large-scale recommendation systems caused by maintaining separate models for different scenarios and objectives, which hinders development velocity and delays technology adoption. To overcome this, the authors propose the Standardized Model Template (SMT) framework, which leverages composable, standardized machine learning components to enable “design once, deploy everywhere,” uniformly accommodating diverse data distributions and optimization objectives. By decoupling model architecture from scenario-specific configurations, SMT reduces the complexity of technology deployment from O(n·2ᵏ) to O(n+k), breaking away from the conventional “one objective, one model” paradigm. Empirical evaluation on Meta’s ad ranking system demonstrates that SMT improves average cross-entropy by 0.63%, reduces engineering time per model iteration by 92%, and increases the throughput of technology-model pair adoption by 6.3×.
This work addresses the disconnect between statistical guarantees and practical performance of conformal prediction (CP) in medical few-shot learning scenarios. We systematically identify and quantify two critical issues—coverage distortion and predictive set inflation—arising from insufficient calibration set size. Through theoretical analysis and empirical evaluation on medical image classification, we apply standard CP frameworks (e.g., quantile regression, inductive CP) and assess performance using both coverage rate and predictive set size. Our results formally demonstrate that, although CP maintains nominal coverage under small calibration sets, predictive intervals become significantly wider, severely degrading practical utility and undermining the precision–reliability trade-off essential for clinical decision-making. This challenges the implicit assumption that “small calibration sets suffice for reliable uncertainty calibration” and provides key empirical evidence and theoretical caution regarding the feasibility boundary of uncertainty quantification in medical AI.
Scientific machine learning experiments often suffer from distorted performance evaluations due to poor experimental design and inconsistent documentation. To address this, we propose a principled framework for ML experimentation tailored to scientific research, encompassing data preprocessing, model selection, cross-validation, and reporting—emphasizing reproducibility, fair comparison, and transparency. Our key contributions include two novel quantitative metrics: the Logarithmic Overfitting Ratio (LOR) and Composite Overfitting Score (COS), which jointly characterize overfitting severity and instability across cross-validation folds. Complementing these, we introduce standardized preprocessing protocols, rigorously defined strong baselines, and modular visualization templates for diagnostic analysis. Empirical evaluation demonstrates that our framework substantially enhances experimental rigor, reproducibility, and result credibility in scientific ML. It further enables robust performance assessment and cross-study comparability, providing systematic support for establishing reliable benchmarks.
This work addresses the lack of distribution-agnostic uncertainty quantification in existing neural operator surrogate models for decision-making tasks, which hinders reliable robust optimization. To overcome this limitation, the study introduces conformal prediction into the function space of neural operator outputs for the first time, leveraging an infinite-dimensional Danskin’s theorem and variational analysis to construct a scalable, robust decision framework that requires no strong distributional assumptions. The proposed method effectively controls regret in downstream tasks and demonstrates significant performance gains over traditional approaches such as Gaussian processes across multiple engineering benchmarks. By providing rigorous theoretical guarantees while enhancing both reliability and computational efficiency, this framework advances the state of the art in robust decision-making under uncertainty.