perturbed-checkpoint selection

Designs and implements procedures that generate and evaluate small perturbations of model checkpoints and measure resulting loss and gradient-pattern changes to select which perturbed checkpoints or which training examples are most likely to produce desired update behavior. Builds extrapolation methods that predict loss and gradient trajectories across fine-tuning from sampled perturbations (loss-pattern extrapolation) to guide checkpoint or example selection.

perturbed-checkpointselection

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.36
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work investigates how to simultaneously enhance functional correctness and computational efficiency in code generation without additional training, by leveraging model weight combinations. The approach trains multiple checkpoints using rewards derived from unit test coverage, then applies linear interpolation and extrapolation of weights, augmented with an inference-time ensemble strategy. The key contribution lies in the first demonstration that weight extrapolation effectively expands the correctness–efficiency Pareto frontier in reinforcement learning for code generation. Experiments show that, on the LCB/hard benchmark, the extrapolation-based ensemble improves the pass@250 metric by 3.3% over the best single model, with consistent gains across three distinct settings: pure inference, tool-augmented reasoning, and agent-based coding.

code reinforcement learningcorrectness-efficiency frontierextrapolative weight averaging

This study addresses the long-standing confounding among validation data volume, checkpoint selection criteria, and final-checkpoint performance in supervised fine-tuning (SFT). By modeling checkpoint selection as a decision problem under limited information and fixing training trajectories, this work decouples these three factors to independently quantify their respective gains. It proposes a selection strategy based on generation accuracy and checkpoint consistency, incorporating bootstrapping for confidence interval analysis. Empirical results demonstrate that expanding the validation budget significantly improves test accuracy, and that generative selection criteria outperform negative log-likelihood (NLL). However, the advantage of such criteria over simply using the final checkpoint remains uncertain. Overall, this research provides a rigorous decoupled analytical framework for SFT evaluation.

checkpoint selectionmodel evaluationsupervised fine-tuning

Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?

Apr 16, 2025
HZ
Hansi Zeng
🏛️ University of Massachusetts Amherst | Google DeepMind | University of Illinois Urbana-Champaign

This work investigates whether pretraining metrics—such as perplexity—reliably predict downstream performance of large language models (LLMs) after fine-tuning, aiming to improve model selection efficiency under fixed computational budgets. The authors formulate checkpoint selection as a pairwise classification task and systematically evaluate 50 distinct 1B-parameter LLM variants across diverse downstream tasks, revealing that perplexity is frequently misleading. They propose novel unsupervised and supervised proxy metrics, which reduce prediction error rates by over 50% in multi-task supervised fine-tuning (SFT) evaluation. This study is the first to empirically demonstrate a non-monotonic relationship between pretraining metrics and downstream performance. The proposed proxies exhibit strong cross-task generalization and practical utility, offering a trustworthy, task-aware evaluation paradigm for optimizing pretraining strategies toward downstream objectives.

Developing new metrics to replace misleading perplexity measuresEvaluating pre-training checkpoints for downstream task performancePredicting fine-tuning outcomes using pre-training indicators

This work addresses the challenge of accurately predicting pretraining loss for large language models across varying model scales, batch sizes, and training steps—particularly under dynamically changing batch sizes and extreme extrapolation of compute budgets. To this end, the authors propose a loss prediction model grounded in a noisy quadratic system, which explicitly models test loss as a function of model size $N$, batch size $B$, and number of weight updates $K$. The framework enables joint optimization of training configurations under composite constraints on time, memory, and computational resources. Notably, it achieves high-precision loss prediction in variable-batch settings—a first in the field—and significantly outperforms existing heuristics such as Chinchilla, even when extrapolating up to 1000× beyond observed compute budgets. The recommended $(N, B, K)$ configurations closely align with empirically optimal solutions.

batch sizecompute budgetlarge language models

Safe Policy Learning through Extrapolation: Application to Pre-trial Risk Assessment

Sep 22, 2021
EB
Eli Ben-Michael
🏛️ Carnegie Mellon University | Harvard Law School | Harvard University | Sun Yat-sen University

Deterministic algorithms used in U.S. pretrial risk assessment lack strategic optimizability and interpretability, limiting their ability to improve upon existing policies safely and transparently. Method: We propose an extrapolation-based safe policy learning framework that integrates robust optimization (maximizing minimum expected utility), partial identification theory, and causal inference—enabling safe policy improvement without assuming randomized interventions. Contribution/Results: To our knowledge, this is the first method to guarantee, with high probability, that a new deterministic policy dominates the current one under observational data, while ensuring statistical safety and interpretability. Evaluated on real-world field experiment data, our approach significantly increases the proportion of defendants in specific subgroups who are safely downgraded to “low risk,” without compromising overall decision safety. This advances responsible deployment of judicial algorithms by providing a principled, auditable, and policy-improving paradigm.

Ensuring safety in policy learning for criminal justiceImproving deterministic pre-trial risk assessment algorithmsMaximizing worst-case expected utility under structural assumptions

Latest Papers

What's happening recently
View more

This work investigates the theoretical underpinnings of memorization and overfitting in stochastic interpolation generative models. Focusing on continuous-time stochastic differential equations and their Euler discretization, it provides the first rigorous theoretical definitions of overfitting and underfitting in generative modeling and derives closed-form expressions for the optimal velocity field and score function. The analysis reveals that generated samples can be expressed as training samples perturbed by three controllable error terms, whose bias is jointly determined by the discretization step size and estimation error. Synthetic experiments corroborate the theoretical prediction that generated samples cluster around the training data distribution, highlighting the critical roles of error accumulation and noise modeling in the model’s reconstruction capability.

estimation errorgenerative modelsmemorization

This work addresses the challenge of reliably generating out-of-distribution data that adheres to new specifications under structural assumptions about the underlying data-generating mechanism. The authors propose a structure-guided extrapolation generation framework that, for the first time, establishes theoretical conditions for approximate identifiability of the target distribution and ensures the reliability of the generated distribution under conservative assumptions. The approach integrates two complementary algorithms—structure-aware optimization and diffusion posterior sampling—to effectively leverage structural priors during generation. Empirical evaluations on both synthetic and real-world image extrapolation tasks demonstrate the framework’s superior performance in terms of generation quality and consistency with the desired extrapolation behavior.

data generating processdistribution identifiabilityextrapolated data generation

This study addresses the challenge of out-of-distribution generalization in scientific machine learning by investigating the theoretical limits of stable extrapolation for high-dimensional functions and operators. Leveraging polynomial, deep neural network, and neural operator frameworks combined with complex analysis and high-dimensional statistics, this work reveals a "blessing of dimensionality" phenomenon: it demonstrates that higher-order coordinate smoothness ensures algebraic error convergence, thereby overcoming the pessimistic bounds of conventional theory. The core contribution lies in deriving explicit convergence rates under arbitrary test measures, establishing extrapolation guarantees that depend solely on the support domain rather than the specific test distribution. Numerical experiments validate the effectiveness of these theoretical findings across extensive extrapolation regimes.

extrapolationhigh-dimensional function learningoperator learning

Hot Scholars

QZ

Quan Zheng

Institute of Software, Chinese Academy of Sciences
Computer Graphics
ZR

Zhiwen Ruan

Southern University of Science and Technology
NLPLLMs
ZC

Zhuoxiao Chen

Oracle | The University of Queensland
3D Scene UnderstandingGeneralization