🤖 AI Summary
This study addresses the vulnerability of large language models to resource consumption attacks caused by uncontrollable repetitive generation, a phenomenon whose underlying mechanisms remain poorly understood. To tackle this issue, this work proposes an intervention mechanism grounded in residual stream dynamics analysis. By monitoring anomalous coordinates within shallow network layers, the approach intercepts the cross-layer propagation of repetitive semantics at its source, thereby efficiently suppressing degenerate behaviors. Empirical evaluations demonstrate that the proposed method reduces cyclic generation rates by 57% on average, effectively mitigating uncontrollable repetition across both multimodal and reasoning models. These findings establish a novel paradigm for enhancing the safety and controllability of large model generation.
📝 Abstract
Uncontrolled repetition can prolong autoregressive generation in large language models (LLMs) and enable resource consumption attacks. Prior analyses of repetitive generation have identified strongly activated features in intermediate and late layers. However, how uncontrolled repetition activity emerges and develops before becoming prominent in these layers remains insufficiently understood. In this paper, we investigate this question primarily in large vision-language models (LVLMs), which support a richer set of uncontrolled repetitions through both visual and textual inputs. We propose Tokenwise Residual Comparison (TRC), a method that identifies and localizes anomalies associated with repetition from residual dynamics during generation. TRC compares attention and multilayer perceptron writes to the residual stream across generated tokens to identify patterns associated with repetition. It then selectively suppresses coordinates in the residual stream at the identified layer. Experiments show that TRC effectively mitigates uncontrolled repetition, reducing loop rates by 57\% on average. Our analysis further shows that repetition semantics emerge in shallow layers and propagate through the residual stream, disrupting normal representations. TRC also generalizes to large language models (LLMs) and large reasoning models (LRMs), where it consistently captures analogous repetition dynamics and achieves effective mitigation. Our work broadens the study of repetitive generation from its prominent internal representations to earlier opportunities for intervention, providing insights for mitigating resource consumption attacks.