🤖 AI Summary
This study addresses the phenomenon whereby large language models generate correct outputs despite corrupted inputs, a process whose internal mechanisms remain poorly understood. By integrating residual stream analysis, attention probing, and fine-tuning techniques, this work reveals a two-stage repair process within the model and identifies an unsupervised spontaneous recovery mechanism. Building upon these findings, we propose a novel paradigm that leverages the hidden states of the first token for efficient fault triage. Experimental results demonstrate that this approach achieves a fault prediction performance with an ROC-AUC of 0.78. Furthermore, moderate fine-tuning significantly enhances model robustness while reducing nonlinear errors.
📝 Abstract
Language models sometimes produce correct outputs even when their inputs are corrupted by deletion, replacement, or misspelling. We study the internal processes accompanying this behavior, which we call context restoration, in controlled attention-only transformers and five pretrained LLMs (1B-32B parameters) across arithmetic, reading comprehension, and multiple-choice reasoning tasks. In the attention-only transformers, restoration emerges spontaneously despite training exclusively on clean sequences, without corruption training or an explicit denoising objective. We find that context restoration follows a two-phase process: early layers localize effects associated with repair at corrupted positions, while later layers accumulate these effects at uncorrupted positions through the residual stream and ultimately concentrate them at the output position. Repair outcome is predictable from hidden states: cosine alignment with the clean state is highly predictive in attention-only models, while linear probes recover additional information in pretrained LLMs. A linear probe using only the corrupted prompt's first-block hidden state predicts failure with mean ROC-AUC 0.78. This enables failure triage under matched or even partially shifted deployment conditions and may reduce unnecessary verification or computation. Failed examples also show substantially greater nonlinearity along corruption directions. Moderate-corruption finetuning increases corruption tolerance while simultaneously reducing displacement-normalized linearization error, associating improved robustness with a more nearly linear response to corruption.