🤖 AI Summary
This study addresses the limitation that a single sentiment direction in large language models cannot capture sentiments modulated by usage factors such as tone and audience. We construct paired controlled datasets to analyze the activation spaces of models including Llama. Through linear probing and residual erasure, we reveal a compact, reproducible residual structure of usage conditions within sentiment representations that is independent of polarity, thereby challenging the conventional linear direction hypothesis. Furthermore, we propose a tone-steering technique applied during generation. Experiments demonstrate that this residual structure significantly influences evaluation metrics, while tone steering achieves a 92.8% preference rate in blind evaluations and simultaneously maintains 98.7% sentiment polarity accuracy.
📝 Abstract
Prior work suggests that sentiment can often be captured by approximately linear directions in LLM activation spaces, but a single direction may not fully capture sentiment representations. In natural communication, sentiment is shaped not only by polarity but also by usage factors, such as tone and audience adaptation. We test whether these factors systematically modulate sentiment representations beyond a shared sentiment direction. We construct a controlled paired dataset that holds event content fixed while varying sentiment polarity and usage factors, and analyze Llama, Mistral, and Gemma. We identify a shared sentiment direction, remove it, and test the residual structure through erasure and generation-time tone steering. Across models, the shared direction is robust (median cosine 0.953-0.975), yet removing it leaves 0.833-0.909 of the original positive-negative representation-difference norm. The residuals contain compact, reproducible usage-conditioned structure. Targeted erasure weakens held-out usage metrics more than random and label-shuffled controls. On Llama, outputs steered along residualized tone components are preferred in 92.8% of blind target-tone comparisons while preserving the requested sentiment polarity in 98.7% of evaluated outputs.