🤖 AI Summary
This study addresses the susceptibility of large language models to hallucinations during reasoning tasks, where existing methods predominantly rely on passive detection or filtering and struggle to intervene at the source. To overcome this limitation, this work proposes USteer, a training-free, inference-time steering mechanism that pioneers the use of uncertainty signals for proactive intervention. By computing confidence gradients to dynamically adjust inter-layer activations, USteer actively reduces output uncertainty without modifying model parameters or requiring additional supervision. This approach facilitates a paradigm shift in hallucination mitigation from passive detection to active control. Experimental results demonstrate that USteer consistently reduces hallucinations across diverse reasoning tasks, validating the feasibility of leveraging confidence signals to directly govern model behavior and enhance generation quality.
📝 Abstract
Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model's confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model's layer-wise activations during inference using the gradient of a confidence measure with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.