Score
Designs and implements procedures that generate and evaluate small perturbations of model checkpoints and measure resulting loss and gradient-pattern changes to select which perturbed checkpoints or which training examples are most likely to produce desired update behavior. Builds extrapolation methods that predict loss and gradient trajectories across fine-tuning from sampled perturbations (loss-pattern extrapolation) to guide checkpoint or example selection.
This work investigates how to simultaneously enhance functional correctness and computational efficiency in code generation without additional training, by leveraging model weight combinations. The approach trains multiple checkpoints using rewards derived from unit test coverage, then applies linear interpolation and extrapolation of weights, augmented with an inference-time ensemble strategy. The key contribution lies in the first demonstration that weight extrapolation effectively expands the correctness–efficiency Pareto frontier in reinforcement learning for code generation. Experiments show that, on the LCB/hard benchmark, the extrapolation-based ensemble improves the pass@250 metric by 3.3% over the best single model, with consistent gains across three distinct settings: pure inference, tool-augmented reasoning, and agent-based coding.
This study addresses the long-standing confounding among validation data volume, checkpoint selection criteria, and final-checkpoint performance in supervised fine-tuning (SFT). By modeling checkpoint selection as a decision problem under limited information and fixing training trajectories, this work decouples these three factors to independently quantify their respective gains. It proposes a selection strategy based on generation accuracy and checkpoint consistency, incorporating bootstrapping for confidence interval analysis. Empirical results demonstrate that expanding the validation budget significantly improves test accuracy, and that generative selection criteria outperform negative log-likelihood (NLL). However, the advantage of such criteria over simply using the final checkpoint remains uncertain. Overall, this research provides a rigorous decoupled analytical framework for SFT evaluation.
This work investigates whether pretraining metrics—such as perplexity—reliably predict downstream performance of large language models (LLMs) after fine-tuning, aiming to improve model selection efficiency under fixed computational budgets. The authors formulate checkpoint selection as a pairwise classification task and systematically evaluate 50 distinct 1B-parameter LLM variants across diverse downstream tasks, revealing that perplexity is frequently misleading. They propose novel unsupervised and supervised proxy metrics, which reduce prediction error rates by over 50% in multi-task supervised fine-tuning (SFT) evaluation. This study is the first to empirically demonstrate a non-monotonic relationship between pretraining metrics and downstream performance. The proposed proxies exhibit strong cross-task generalization and practical utility, offering a trustworthy, task-aware evaluation paradigm for optimizing pretraining strategies toward downstream objectives.
This work addresses the challenge of accurately predicting pretraining loss for large language models across varying model scales, batch sizes, and training steps—particularly under dynamically changing batch sizes and extreme extrapolation of compute budgets. To this end, the authors propose a loss prediction model grounded in a noisy quadratic system, which explicitly models test loss as a function of model size $N$, batch size $B$, and number of weight updates $K$. The framework enables joint optimization of training configurations under composite constraints on time, memory, and computational resources. Notably, it achieves high-precision loss prediction in variable-batch settings—a first in the field—and significantly outperforms existing heuristics such as Chinchilla, even when extrapolating up to 1000× beyond observed compute budgets. The recommended $(N, B, K)$ configurations closely align with empirically optimal solutions.
Deterministic algorithms used in U.S. pretrial risk assessment lack strategic optimizability and interpretability, limiting their ability to improve upon existing policies safely and transparently. Method: We propose an extrapolation-based safe policy learning framework that integrates robust optimization (maximizing minimum expected utility), partial identification theory, and causal inference—enabling safe policy improvement without assuming randomized interventions. Contribution/Results: To our knowledge, this is the first method to guarantee, with high probability, that a new deterministic policy dominates the current one under observational data, while ensuring statistical safety and interpretability. Evaluated on real-world field experiment data, our approach significantly increases the proportion of defendants in specific subgroups who are safely downgraded to “low risk,” without compromising overall decision safety. This advances responsible deployment of judicial algorithms by providing a principled, auditable, and policy-improving paradigm.
This work investigates the theoretical underpinnings of memorization and overfitting in stochastic interpolation generative models. Focusing on continuous-time stochastic differential equations and their Euler discretization, it provides the first rigorous theoretical definitions of overfitting and underfitting in generative modeling and derives closed-form expressions for the optimal velocity field and score function. The analysis reveals that generated samples can be expressed as training samples perturbed by three controllable error terms, whose bias is jointly determined by the discretization step size and estimation error. Synthetic experiments corroborate the theoretical prediction that generated samples cluster around the training data distribution, highlighting the critical roles of error accumulation and noise modeling in the model’s reconstruction capability.
本文提出了一种结合物理指导的机器学习框架,通过BiLSTM和PINN解决工程中模型外推问题,并在经典瞬态扩散问题上验证了其有效性。
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
This work addresses the challenge of reliably generating out-of-distribution data that adheres to new specifications under structural assumptions about the underlying data-generating mechanism. The authors propose a structure-guided extrapolation generation framework that, for the first time, establishes theoretical conditions for approximate identifiability of the target distribution and ensures the reliability of the generated distribution under conservative assumptions. The approach integrates two complementary algorithms—structure-aware optimization and diffusion posterior sampling—to effectively leverage structural priors during generation. Empirical evaluations on both synthetic and real-world image extrapolation tasks demonstrate the framework’s superior performance in terms of generation quality and consistency with the desired extrapolation behavior.
This study addresses the challenge of out-of-distribution generalization in scientific machine learning by investigating the theoretical limits of stable extrapolation for high-dimensional functions and operators. Leveraging polynomial, deep neural network, and neural operator frameworks combined with complex analysis and high-dimensional statistics, this work reveals a "blessing of dimensionality" phenomenon: it demonstrates that higher-order coordinate smoothness ensures algebraic error convergence, thereby overcoming the pessimistic bounds of conventional theory. The core contribution lies in deriving explicit convergence rates under arbitrary test measures, establishing extrapolation guarantees that depend solely on the support domain rather than the specific test distribution. Numerical experiments validate the effectiveness of these theoretical findings across extensive extrapolation regimes.