Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing

📅 2026-07-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates how systematic variations in linguistic constructions influence large language models’ (LLMs’) judgments of political stance, introducing constructional alternation into the analysis of LLM stance stability for the first time. By generating six types of controlled rewrites that either preserve or invert original meaning and applying causal interventions via activation patching across four open-source LLMs, the authors trace the internal causal pathways underlying stance attribution. Their experiments reveal that mid-to-late decoder layers—particularly block outputs at the final prompt position—contribute most significantly to recovering the original stance distribution. This finding indicates that sensitivity to political stance is localized within specific model components, offering causal evidence for understanding the social-cognitive mechanisms of LLMs.
📝 Abstract
Large language models (LLMs) are known to be sensitive to prompt and input formulations. However, existing studies have focused on lexical realization and largely ignored constructional choice. This paper studies whether linguistic construction can systematically shift LLM decisions and where these shifts can be causally localized inside the model. We use political stance judgment as a meaning-sensitive case study and extend an English political statements dataset, resulting in six controlled linguistic rewrite types that preserve or invert the meaning of a statement. Experiments on four open-weight models show that stance instability affect both meaning-preserving and meaning-inversing rewrites. Because output shifts reveal that rewrites affect stance, but not where in the model, we apply activation patching, where activations from the original statement are substituted into the forward pass for the rewritten statement and measure which components recover the original stance distribution. The results show that mid-to-late decoder layers, especially block outputs at the final prompt position, provide the strongest restoration signal.
Problem

Research questions and friction points this paper is trying to address.

linguistic realization
large language models
stance detection
constructional choice
causal tracing
Innovation

Methods, ideas, or system contributions that make the work stand out.

constructional choice
causal tracing
activation patching
stance instability
linguistic realization
🔎 Similar Papers
No similar papers found.