🤖 AI Summary
This study investigates whether large language models can perceive the degradation of their computational substrate caused by quantization. Inspired by agnosia, we explore the self-monitoring capabilities of models using linear probes and LoRA fine-tuning to extract and analyze “quantization fingerprints” from both internal representations and generated text. We demonstrate that internal representations offer greater interpretability than external outputs and propose a quantization-aware pathway based on these internal fingerprints. Our findings confirm that while models struggle to identify their quantization state from isolated outputs, quantization fingerprints can be effectively decoded from internal representations. Furthermore, this work reveals limitations in the generalizability of such perception mechanisms across different quantization methods, offering novel insights into the underlying self-awareness of large language models.
📝 Abstract
Can LLMs recognize degradation in their own computational substrate? Inspired by anosognosia, a neurological condition in which patients fail to recognize impairments in their own abilities, we investigate whether LLMs can recognize degradation in their computational substrate induced by quantization. We first show that existing models fail to self-report their quantization state, even when provided with their own generated text as an external cue. Linear probing reveals that, while generated text carries almost no trace of quantization, internal representations contain clear, method-specific fingerprints. Through training, models learn to identify severely degraded outputs such as those of 4-bit models by comparison, yet still fail to do so from a single output. A shared LoRA trained jointly across quantization levels succeeded in reading out internal fingerprints, but fails on unseen quantization methods, merely mapping method-specific fingerprints to labels. Whereas external self-observation can restore awareness in some cases of human anosognosia, our results suggest that the more promising route to enabling such awareness in LLMs may lie in their internal representations. Our results highlight fundamental limits of generalizability to LLM self-monitoring.