Score
A space-filling experimental-design method that generates representative sample points across multi-dimensional parameter spaces to enable efficient surrogate-model training and to study interactions with estimators and resampling methods like the jackknife.
This paper addresses the challenge of constructing space-filling designs for computer experiments under high-dimensional and constrained settings, aiming to enhance prediction accuracy and uncertainty quantification—particularly for Gaussian process (GP) surrogates. We systematically analyze the theoretical connections between space-filling criteria (e.g., Maximin, Latin hypercube, projection-based designs) and GP performance, establishing—for the first time—a rigorous link between fill distance and predictive error bounds. We further propose failure criteria for design construction in high-dimensional and constrained scenarios and identify the convergence of adaptive sampling with machine learning as a key evolutionary direction. Through comprehensive numerical experiments, we quantitatively evaluate trade-offs among accuracy, robustness, and computational cost across design families. Results confirm that fill distance serves as a strong indicator of surrogate generalization capability. The work delivers interpretable, reusable design principles for digital twin and cyber-physical systems.
This study addresses the challenge that existing factor screening designs struggle to simultaneously achieve high screening efficiency and adequate space-filling properties, thereby limiting the accuracy of subsequent surrogate modeling. To overcome this limitation, the authors propose a novel class of one-factor-at-a-time (OFAT) designs that systematically incorporates space-filling characteristics into the OFAT framework for the first time. While preserving the inherent efficiency of OFAT in identifying active factors, the proposed approach significantly enhances coverage of the input space. By refining the MOFAT family of designs through optimization-based space-filling criteria, the method demonstrates superior performance in both factor identification accuracy and space-filling quality across multiple numerical experiments, effectively balancing the dual objectives of efficient screening and high-fidelity modeling.
Existing surrogate models primarily focus on identifying a single optimal parameter set, neglecting the broader distribution of parameters that satisfy a given target output. Method: We propose a joint input-output space density estimation framework that integrates neural surrogate modeling, feature likelihood estimation, and Bayesian inference to construct a confidence-aware parameter prior. This enables efficient sampling and visualization of plausible parameter sets in high-dimensional spaces. Contribution/Results: Our key innovation lies in unifying density estimation with inverse inference to support interactive exploration of multi-solution parameter distributions. Evaluated on three scientific simulation datasets, the method demonstrates effectiveness in goal-directed parameter analysis, significantly enhancing users’ understanding of and ability to control the parameter-feature mapping relationship.
This work addresses the sampling bias in traditional Maximum Projection (MaxPro) designs caused by the use of Euclidean distance over bounded domains, which leads to under-sampling near boundaries—particularly corners—and introduces bias in Monte Carlo estimates. To remedy this, the authors propose the uniform Maximum Projection (uMaxPro) method, which replaces Euclidean distance with a periodic distance based on the minimum image convention. This modification preserves the Latin hypercube structure while eliminating boundary-induced density distortions, thereby achieving both uniformity in low-dimensional subspaces and overall statistical uniformity. The uMaxPro design significantly reduces bias and variance in Monte Carlo integration, enhances subspace projection performance and discrepancy, and improves the accuracy of surrogate models and the reliability of probabilistic estimates in engineering benchmark problems such as mesoscale finite element modeling of concrete.
For computer experiments involving variables with a grouped additive structure—i.e., no interactions between groups—this paper proposes Group Orthogonal Arrays (GOAs) as a novel design paradigm surpassing conventional space-filling approaches. We establish, for the first time, a systematic theoretical framework for orthogonal arrays tailored to grouped additive models, supporting arbitrary prime-power factor levels and flexible run sizes. Integrating finite-field algebra, combinatorial design theory, and orthogonal array construction techniques, we develop multiple explicit construction algorithms. These yield large-scale, practical GOA tables, substantially expanding the scope of feasible experimental designs. Empirical evaluation demonstrates that GOAs achieve superior within-group projection uniformity compared to state-of-the-art methods and reduce average prediction error in response surface modeling by 12%–28%.
This work addresses the construction of space-filling designs for computer experiments by proposing two quasi-Monte Carlo (QMC) lattice-based algorithms: one based on rank-1 lattices and the other on Korobov lattices. Uniformity is characterized through a unified framework involving covering and separation radii, enabling the first explicit construction of lattice point sets that approximate quasi-uniform Kronecker sequences. The authors prove that these constructions achieve an isotropic discrepancy of optimal theoretical order $O(N^{-1/d})$. Furthermore, they enhance isotropy by optimizing the Korobov generating vector via the LLL lattice basis reduction algorithm. Numerical experiments demonstrate that the proposed designs outperform existing QMC point sets in terms of quasi-uniformity and significantly improve predictive accuracy in Gaussian process regression tasks.
This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.
This study addresses key limitations in existing sequential experimental design strategies for Kriging models, including low information utilization in single-point approaches, insufficient research on batch strategies, and the tendency of sampling points to cluster. To overcome these issues, the authors propose two novel single-point sequential criteria and develop a general-purpose batch sequential design framework that, for the first time, enables efficient extension of arbitrary sequential criteria to batch settings. The framework effectively mitigates point clustering and substantially enhances both information utilization efficiency and experimental cycle effectiveness. Numerical experiments based on Kriging surrogate models demonstrate that the proposed methods consistently outperform state-of-the-art approaches across multiple benchmark functions, achieving superior approximation accuracy and computational efficiency.
This study addresses the challenge of identifying valid short-term surrogate endpoints when long-term outcomes are costly or infeasible to observe in randomized experiments. Existing causal criteria for surrogacy are often non-identifiable, limiting their practical utility. To overcome this, the authors propose a plug-in composite surrogate learning framework that directly optimizes the predictive performance of treatment effect estimation, thereby circumventing traditional identifiability constraints. The approach constructs surrogates from post-treatment variables and introduces two learning strategies explicitly tailored to effect prediction. Theoretical analysis demonstrates that the method yields unbiased effect estimates under standard assumptions, and empirical evaluations on both synthetic and real-world datasets show that the learned surrogates significantly outperform existing approaches in predicting the primary treatment effect.
Data-driven polynomial chaos expansion (PCE) surrogate models lack reliable predictive uncertainty quantification. Method: This paper proposes a tightly integrated interval estimation framework combining jackknife conformal prediction with PCE. Leveraging the linear structure of PCE regression, we derive closed-form analytical expressions for leave-one-out residuals and predictions—eliminating the need for repeated model retraining or an independent calibration dataset. Contribution/Results: The method achieves both computational efficiency and data economy. Extensive validation on benchmark problems confirms the statistical validity and coverage accuracy of the resulting prediction intervals. Furthermore, it quantitatively characterizes how training sample size influences interval width and reliability. Crucially, this work establishes the first retraining-free, theoretically guaranteed, plug-and-play uncertainty quantification framework for PCE—providing rigorous, distribution-free confidence intervals without sacrificing predictive fidelity.