data contamination modeling

Designs and analyzes probabilistic and simulation models that represent processes by which datasets become contaminated and that generate contaminated-data scenarios, and builds tools to propagate those scenarios through algorithms and estimators. Uses those models to quantify bias, variance, and unbounded or worst‑case failure modes introduced by contamination and to estimate the statistical behavior and performance of methods under contaminated distributions.

datacontaminationmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

What Quality Engineers Need to Know about Degradation Models?

Jul 19, 2025
JM
Jared M. Clark
🏛️ Virginia Tech | University of South Florida | Brigham Young University | SAS

Quality engineers lack systematic degradation modeling methodologies, hindering the accuracy and practical implementation of reliability assessment. Method: This study establishes an industrially oriented degradation analysis framework that unifies diverse degradation data sources—including repeated measurements and accelerated destructive testing—and integrates path models (e.g., general path models) with stochastic process models (e.g., Wiener processes), augmented by Bayesian and likelihood-based statistical inference techniques. A standardized modeling workflow and lifetime prediction toolkit are implemented in R/Python. Contribution/Results: The framework bridges the gap between theoretical degradation modeling and engineering practice, significantly improving the accuracy and reproducibility of reliability predictions for complex systems. It delivers an actionable guideline and open-source software support for industry-standardized deployment, enabling robust, traceable, and scalable reliability engineering.

Address reliability estimation and prediction in diverse industriesIntroduce degradation models for quality engineering applicationsReview data types and modeling approaches for degradation

Machine learning models frequently suffer unexpected failures in real-world deployment, hindering practical adoption. Method: This paper introduces, for the first time, an orthogonal dichotomy framework distinguishing reliability from robustness, formally characterizing model failure mechanisms from first principles and systematically mapping them to engineering practices and real-world deployment scenarios. Our approach integrates probabilistic modeling, uncertainty quantification, adversarial robustness analysis, distributional shift detection, and system-level fault tree analysis—bridging theoretical insights with industrial-grade diagnostic tools and canonical failure case studies. Contribution/Results: We deliver an actionable failure attribution guide comprising rigorous theoretical foundations, an open-source toolchain, and cross-domain application exemplars. The framework significantly enhances model trustworthiness, debuggability, and deployment success rates.

Differentiating failure causes: reliability vs robustnessProviding practical techniques for ML model reliabilityUnderstanding unexpected failures in ML models

Simulations in Statistical Workflows

Mar 31, 2025
PB
Paul-Christian Burkner
🏛️ TU Dortmund University | Independent Scientist | Rensselaer Polytechnic Institute

This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.

Analyzing trends in simulation-based statistical algorithmsExamining simulation roles in statistical workflowsExploring future impacts of simulations on statistics

The pushed beta distribution and contaminated binary sampling

Mar 14, 2025
BO
Ben O'Neill
🏛️ Allen Consulting

This paper addresses the limited modeling capability of the standard Beta distribution in contaminated binary sampling. We propose and systematically study the “Pushed Beta Distribution”—a generalization of the Gauss hypergeometric Beta distribution, obtained by introducing a directional multiplicative term into its density kernel. First formally defined herein, this distribution is rigorously proven to be the exact conjugate posterior for contaminated binary models. We fully characterize its analytical properties—including moments, mode, and asymptotic behavior—and establish an efficient computational framework based on numerical integration and special functions. Furthermore, we develop practical algorithms for computing the cumulative distribution function, quantiles, and generating random samples. These contributions extend the theoretical boundaries of the Beta family and provide a new Bayesian inference tool for contaminated binary data that balances statistical interpretability with computational tractability.

Analyzing Gaussian hypergeometric beta distribution variantsDeveloping computational methods for directional distribution propertiesModeling contaminated binary sampling using Bayesian inference

Two-stage Design for Failure Probability Estimation with Gaussian Process Surrogates

Oct 06, 2024
AS
Annie S. Booth
🏛️ Virginia Tech | Penn State

This work addresses the challenge of estimating small failure probabilities under stochastic inputs in computationally expensive deterministic simulations. We propose a two-stage adaptive budget allocation framework: in Stage I, a Gaussian process surrogate is sequentially trained using a contour-localization strategy; in Stage II, remaining simulation budget is greedily allocated to critical regions—guided by classification entropy—to perform high-fidelity evaluations. A hybrid Monte Carlo estimator is then constructed by integrating surrogate predictions with observed high-fidelity responses. Our method introduces the first “exploration–exploitation decoupled” budget allocation paradigm, overcoming reliability limitations inherent in pure surrogate-based Monte Carlo and importance sampling. Experiments across multiple benchmark functions and an airfoil flow simulation demonstrate that the approach achieves significantly improved accuracy and robustness using only several hundred high-fidelity evaluations.

Estimating failure probabilities with limited computational budgetImproving efficiency over existing sequential contour location methodsOptimizing surrogate model training for accurate classification

Latest Papers

What's happening recently
View more

This work proposes the first general framework to systematically quantify and apportion epistemic uncertainty arising from substituting true subprocesses with approximate or learned submodels in stochastic simulation and digital twin applications. The framework constructs confidence or credible intervals for performance metrics via bootstrapping and Bayesian model averaging, and employs a tree-based decomposition to allocate total output variability to individual submodels, yielding importance scores. It is compatible with both parametric and nonparametric models, supports frequentist and Bayesian paradigms, and accommodates dynamic initialization scenarios. Validation on synthetic data and a call center digital twin demonstrates that the method effectively reveals each submodel’s contribution to overall uncertainty, significantly enhancing the interpretability and reliability of simulation outcomes.

digital twinsepistemic uncertaintyoutput variability

Conservative Software Reliability Assessments Using Collections of Bayesian Inference Problems

Nov 10, 2025
KS
Kizito Salako
🏛️ The Centre for Software Reliability | Department of Computer Science | City St. George's, University of London

To address conservative reliability assessment for safety-critical software under prior uncertainty, this paper proposes a robust Bayesian framework that computes the worst-case posterior predictive probability of fault-free operation—thereby yielding a conservative estimate of future reliability. Methodologically, software failures are modeled as a Bernoulli process, and the approach integrates set-based Bayesian inference with asymptotic analysis. Key contributions include: (1) the first closed-form analytical solution for the worst-case posterior predictive probability; (2) characterization of its asymptotic convergence properties; and (3) an extension of robust Bayesian theory, providing a rigorous mathematical foundation for quantifying worst-case behavior under prior uncertainty. The framework balances theoretical rigor with practical applicability, enabling high-assurance software reliability certification.

Determining worst-case posterior predictive probabilities for software reliabilityExtending robust Bayesian methods for safety-critical software assessmentsModeling software failures using Bernoulli process Bayesian inference

Formal Analysis of Metastable Failures in Software Systems

Oct 03, 2025
RI
Rebecca Isaacs
🏛️ AWS | UC Santa Cruz | MPI-SWS | University of Birmingham

Metastable failures are rare yet high-risk failure modes in cloud systems, triggered by transient load spikes and persisting as prolonged performance degradation even after stress subsides. This paper addresses request-response server systems by proposing a continuous-time Markov chain (CTMC)-based modeling framework. It introduces the first formal definition of metastability via escape probability and establishes a quantitative relationship between metastability and the spectral gap of the CTMC’s dominant eigenvalues—enabling computationally tractable recovery-time prediction and visual identification of metastability-prone parameter configurations. The methodology integrates domain-specific language modeling, data-driven calibration, and combined qualitative/quantitative analysis. The developed tool detects diverse real-world metastable phenomena within milliseconds. Experimental validation confirms the critical phenomenon: as system parameters approach the metastable regime, recovery time grows exponentially.

Analyzing metastable failures in large-scale software systemsModeling server systems using continuous-time Markov chainsPredicting recovery times and identifying metastable parameterizations

Hot Scholars

DM

Debmalya Mandal

Assistant Professor, University of Warwick
Computational Social ChoiceAlgorithmic FairnessReinforcement Learning
MC

Michele Caprio

Lecturer (Asst. Prof.), The University of Manchester
Imprecise ProbabilityApplied ProbabilityArtificial IntelligenceStatistical Theory
AP

Antonio Punzo

Full Professor of Statistics, University of Catania
Mixture ModelsHidden Markov ModelsHeavy-tailed DistributionsSerial Dependence
AZ

Andreas Züfle

Associate Professor at Emory University
Spatial DatabasesData MiningGeosimulationSpatial Computing