Learning Perturbations to Extrapolate Your LLM

πŸ“… 2026-05-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the limited out-of-distribution extrapolation capability of current large language models, which stems from their reliance on fixed discrete perturbations that lack flexibility. To overcome this limitation, the authors propose a continuous, learnable perturbation framework that introduces trainable latent vectors in the embedding space to flexibly perturb token prefixes. The method integrates unbiased estimating equations with stochastic gradient descent for optimization and provides a theoretical analysis of the estimator’s statistical properties under over-parameterization theory. Experimental results demonstrate that the proposed approach significantly outperforms multiple state-of-the-art baselines on both synthetic and real-world datasets, achieving robust performance gains in out-of-distribution settings.
πŸ“ Abstract
Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on discrete perturbations with fixed designs, which limits their flexibility. In this work, we propose a framework where token prefixes are perturbed by a learnable transformation of a continuous latent vector within an embedding space. To overcome the challenge of an intractable marginal likelihood, we derive unbiased estimating equations for model parameters and optimize them via stochastic gradient descent. We establish the statistical properties of the resulting estimator in over-parameterized regimes. Empirical evaluations on both synthetic and real-world datasets demonstrate that our proposal yields significant gains in out-of-domain settings over a range of state-of-the-art baseline methods.
Problem

Research questions and friction points this paper is trying to address.

extrapolation
large language models
perturbations
out-of-domain generalization
flexibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

learnable perturbations
continuous latent vector
extrapolation
large language models
unbiased estimating equations
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
Z
Zetai Cen
School of Mathematics, University of Bristol
C
Chenfei Gu
School of Statistics and Data Science, Shanghai University of Finance and Economics
Jin Zhu
Jin Zhu
School of Mathematics, University of Birmingham
machine learning
T
Ting Li
School of Statistics and Data Science, Shanghai University of Finance and Economics
Yunxiao Chen
Yunxiao Chen
Department of Statistics, London School of Economics and Political Science
Multivariate StatisticsPsychometrics
Chengchun Shi
Chengchun Shi
London School of Economics and Political Science
Large Language ModelsReinforcement LearningStatistics