Score
Designs and implements procedures and software that apply rejection sampling and output filtering to model outputs after generation: defining acceptance criteria and confidence thresholds, running resample-or-aggregate loops that discard or reweight outputs that fail constraints, and combining filtering with aggregation methods. Builds and analyzes output-level quality-control mechanisms including statistical acceptance tests, thresholding rules, and conditional (e.g., rotation-aware) filters to ensure retained outputs meet specified properties.
ISO 2859-2 exhibits inaccurate estimation of residual lot quality risk—particularly for small lots—and inflated consumer’s risk in destructive testing due to non-recoverable samples. To address this, this paper proposes a novel Bayesian attribute acceptance sampling method. It innovatively treats the residual lot size as a fixed parameter and integrates a hypergeometric likelihood with a reference prior to construct a decision-theoretic framework that rigorously controls consumer’s risk. The resulting standardized sampling plans are compact and computationally efficient, substantially reducing actual consumer’s risk in small-lot scenarios. This approach extends the applicability of acceptance sampling beyond the limitations of conventional standards in destructive inspection settings. Its theoretical foundation—grounded in objective Bayesian inference—and empirical efficacy support its potential adoption into international standardization frameworks.
In A/B testing, rigorously evaluating novel estimation algorithms—when the true treatment effect is unobserved—remains a fundamental methodological challenge. This paper establishes, for the first time, a comprehensive theoretical framework for estimation and inference based on sample splitting: it derives the asymptotic distribution of sample-split estimators and characterizes their bias structure relative to full-sample performance; introduces a bias–variance trade-off analytical paradigm and proposes a correction-based confidence interval construction method. Leveraging statistical inference, asymptotic theory, Monte Carlo simulation, and empirical validation, the framework enables robust, production-grade evaluation of new algorithms within industrial A/B testing platforms. Theoretical results are thoroughly validated via simulation studies. The proposed infrastructure enhances A/B testing methodology by delivering an interpretable, reproducible, and deployable evaluation system.
This paper addresses three classical hypothesis testing problems in high-dimensional settings: (1) testing mean vector differences between correlated or independent multivariate samples, (2) testing whether a mean vector equals a specified constant vector, and (3) goodness-of-fit testing for a prescribed distribution. We propose a general nonparametric testing framework based on rejection sampling. Unlike conventional methods, it avoids asymptotic distributional assumptions and instead constructs the exact finite-sample distribution of the test statistic via Monte Carlo simulation coupled with an accept-reject mechanism. Its key innovation lies in the first systematic integration of rejection sampling into statistical test construction—yielding dimension-agnostic performance, implementation simplicity, and high statistical power. Simulation studies demonstrate that the method approaches the uniformly most powerful test in mean vector testing and achieves superior performance in distributional goodness-of-fit testing. Overall, it establishes a scalable, robust, and reproducible paradigm for high-dimensional nonparametric inference.
This study addresses the challenge of false positive accumulation caused by verifiers in adaptive agent generation, as well as the absence of reliable stopping criteria within generate-verify loops. To overcome these limitations, this work proposes an e-value analytical framework based on exponential betting, coupled with a novel conformal risk control procedure tailored for non-monotonic losses. By integrating distribution-free statistical testing, the proposed theory rigorously bounds the false discovery rate (FDR) of accepted proposals. Experiments conducted on both synthetic scenarios and protein design benchmarks validate the effectiveness of the approach. Ultimately, this research provides a statistically guaranteed adaptive termination mechanism for agent workflows, ensuring robust and reliable generation processes without compromising theoretical safety guarantees.
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
This study addresses the sensitivity to noise schedules and the lack of theoretical justification in multi-step sampling for consistency models. By analyzing the composition of noising and denoising operators, it establishes a non-asymptotic convergence theory under explicitly verifiable stability assumptions. Methodologically, the analysis decouples initialization error contraction from approximation error accumulation, revealing that large early-stage noise drives contraction while small late-stage noise controls residual bias, with explicit constants derived for strongly log-concave targets. Experiments confirm that the theoretically predicted contraction and approximation profiles are reliably measurable. This work provides both rigorous theoretical guidance and a practical framework for designing multi-step consistency samplers.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
This study addresses the challenge of contaminated SCADA data from wind turbines, which often includes anomalies, transients, and non-steady-state operating points, rendering traditional expert-driven manual filtering inefficient. To overcome this limitation, the authors propose an unsupervised filtering approach that leverages multivariate feature engineering combined with multiple clustering algorithms to automatically detect both explicit and implicit abnormal operating conditions. A key innovation lies in the design of robust evaluation metrics tailored for unlabeled SCADA data, moving beyond the conventional reliance on power curves alone. Experimental results demonstrate that the proposed method consistently outperforms manual filtering across most scenarios, preserving a higher proportion of valid operational data while requiring only minimal expert calibration for deployment.
This study addresses the challenge of effectively validating input model specifications in digital twin simulations, where conventional approaches—relying solely on marginal output distributions—often fail to detect misspecified joint input models. To overcome this limitation, the authors propose a novel statistical validation framework based on sub-trajectory conditioning. By repeatedly restarting simulations from observed system states while conditioning on subsets of random inputs, the method constructs conditional output distributions that enable goodness-of-fit testing of the full joint input model. This approach innovatively transcends the constraints of marginal validation and is complemented by diagnostic tools to pinpoint specific input sources responsible for detected discrepancies. Empirical evaluations on M/M/1 and tandem queueing systems demonstrate the framework’s heightened sensitivity and effectiveness, successfully identifying input model misspecifications that traditional methods overlook.
研究通过引入双前缀框架,解决了预执行监督中验证单元选择问题,发现较短的验证单元能提高零样本监控器的判别能力。