Score
Designs and implements decision procedures and predictive models that monitor iterative training or evaluation trajectories to decide when to halt or continue computation. This includes classifier-based and learned stopping rules, per-instance or per-task adaptive stopping points, convergence and overfitting diagnosis, selection of early-stopping criteria and strategies, and analysis of calibration and formal guarantees.
In scientific computing applications—such as trajectory prediction, optimal control, and minimum energy path computation—downstream algorithms critically depend on accurate model evaluations. Conventional mean-squared-error-based supervised learning often induces task-specific performance degradation due to misalignment between the loss function and the ultimate algorithmic objective. Method: We propose a task-oriented predictive modeling paradigm that replaces standard regression losses with a surrogate objective: the maximum prediction error over a downstream task support set. Our framework integrates sampling measure modeling, empirical risk discretization, and iterative optimization to directly optimize downstream algorithmic performance. Contribution/Results: This is the first approach to explicitly embed downstream robustness requirements into the training objective. Evaluated across multiple scientific computing benchmarks, it consistently improves both predictive accuracy and algorithmic stability, demonstrating superior generalization under task-relevant perturbations.
A long-standing debate in reinforcement learning for large language models concerns the relative statistical difficulty of process supervision versus outcome supervision, with conventional wisdom favoring the former. Method: This paper theoretically analyzes their statistical complexity under standard data coverage assumptions and introduces the trajectory measure transformation lemma—a novel tool that formally links return-oriented trajectory distributions to step-level distributional shifts. It further proves that any policy’s advantage function serves as an optimal process reward model. Results: The work establishes statistical equivalence between process and outcome supervision, demonstrating that observed performance gaps stem from algorithmic implementation flaws—not intrinsic statistical hardness. By unifying the theoretical foundations of both paradigms, it provides rigorous justification for lightweight, high-efficiency outcome supervision, challenging the prevailing reliance on fine-grained process annotations.
This paper addresses the lack of a unified early-stopping mechanism for implicit regularization in iterative learning. To bridge this gap, the authors propose a theory-driven early-stopping framework. They develop EarlyStopping, an open-source Python toolkit that—uniquely—systematically integrates truncated SVD, Landweber iteration, conjugate gradient, L2-boosting, and regression trees. The toolkit supports user-defined data generation and enables real-time monitoring of theoretical regularization strength, including effective degrees of freedom and bias–variance trade-offs. Implemented in NumPy/SciPy, it provides sequential risk estimation and analytically derived stopping boundaries. Experiments reproduce key theoretical results on implicit regularization, demonstrating that principled early stopping effectively suppresses noise propagation, constrains generalization error growth, and significantly enhances algorithmic robustness and interpretability—thereby narrowing the gap between theoretical analysis and practical deployment.
Large language models (LLMs) exhibit insufficient reasoning capabilities for complex programming tasks: process supervision relies on costly and error-prone reward modeling, while outcome supervision struggles to coordinate multi-step reasoning. To address this, we propose a novel “outcome-refinement-as-process” supervision paradigm that eliminates explicit reward modeling and instead leverages program execution feedback—such as runtime outputs and error traces—as label-free, reliable intermediate supervision signals. Our approach integrates tree-based multi-path exploration with a lightweight model adaptation framework to enable efficient, execution-guided reasoning. Evaluated across five LLMs and three benchmark datasets, our method achieves average improvements of 26.9% in code correctness and 42.2% in execution efficiency. Notably, it significantly boosts the performance of smaller models on algorithmic competition–style tasks. This work establishes a scalable, low-overhead paradigm for complex programming reasoning, grounded in direct execution feedback rather than surrogate reward signals.
In Bayesian sequential trials, error rate evaluation relies on computationally expensive Monte Carlo simulations, hindering efficient optimization of sample size and decision thresholds. Method: This paper establishes, for the first time, analytical functional relationships between posterior and posterior predictive probabilities and sample size. Leveraging Bayesian decision theory and asymptotic analysis—combined with numerical fitting and error-rate inversion—the method enables precise error-rate assessment for any sample size using only two simulations, and rapidly identifies optimal design parameters. Contribution/Results: The approach drastically reduces computational cost while achieving error-rate control accuracy comparable to conventional simulation-based methods. In two real-world case studies, it attains exact error-rate calibration and accelerates design optimization by several orders of magnitude. This provides a scalable, verifiable, and highly efficient design paradigm for Bayesian adaptive trials.
This work addresses the challenge of achieving optimal performance in reasoning models under minimal computational cost by proposing LearnStop, a lightweight, learning-based early-stopping mechanism that operates without access to hidden states. LearnStop dynamically predicts whether the current reasoning prefix is correct by online aggregation of multidimensional features—such as confidence, entropy, and answer stability—at fixed budget points to decide whether to terminate inference early. Theoretical analysis and experiments demonstrate that LearnStop significantly outperforms conventional threshold-based methods specifically in tasks lacking reliable single-scalar signals yet containing early correct answers, such as open-ended mathematical reasoning; on GSM8K, it improves accuracy by 2.8% over the strongest scalar baseline and extends the performance frontier under fixed budgets. However, its advantage is limited in multiple-choice or extremely difficult problems. The study systematically delineates the effectiveness boundary of learned stopping strategies.
This study investigates the impact of model selection criteria—such as accuracy versus loss—on test performance in neural classifier training, particularly under early stopping with patience. Through systematic empirical evaluation using k-fold cross-validation on standard benchmarks, the work compares multiple validation metrics, including accuracy, cross-entropy, C-Loss, and PolyLoss, under both early stopping and post-hoc full-trajectory selection strategies. The findings reveal that validation loss–based criteria consistently outperform validation accuracy, which exhibits not only inferior performance but also lower stability. More critically, regardless of the selection criterion employed, the chosen models are typically substantially worse than the best test performance observed during training, thereby exposing a fundamental limitation in current model selection paradigms.
This study addresses the suboptimal decision-making in document screening, where existing stopping strategies focus solely on recall while neglecting the actual costs and benefits of review tasks. To bridge this gap, the work introduces decision theory into the problem for the first time, deriving three adaptive stopping strategies grounded in the Expected Value of Perfect Information (EVPI). These strategies are integrated within a Technology-Assisted Review (TAR) framework to align stopping decisions with task-specific utility objectives. Empirical evaluations on the CLEF-IP patent dataset and medical systematic review corpora demonstrate that the proposed approach significantly improves net utility compared to prevailing stopping rules, achieving better alignment between screening outcomes and real-world review goals.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.