Auditing Alignment Controllability in LLMs via Political Axes

πŸ“… 2026-07-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study challenges the conventional practice of treating large language models as fixed points on political spectra, which overlooks the dynamic influence of system prompts on their outputs. By conducting large-scale prompt manipulation experiments across seven mainstream models, twelve ideological personas, and seventy Political Compass questions, the authors propose a controllability auditing framework centered on dispersion, symmetry, saturation, and refusal thresholds. Leveraging multi-role prompt engineering, variance analysis, and geometric displacement modeling on 63,700 model responses, they quantify the adjustable range of political stance. Results reveal that contextual prompts account for 88%–93% of variance along political axes, while model identity contributes less than 3%. Models exhibit divergent saturation behaviors under extreme prompts yet converge toward authoritarian-leaning shifts in authoritarian contexts, resolving prior contradictions stemming from non-centered baselines.
πŸ“ Abstract
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.
Problem

Research questions and friction points this paper is trying to address.

alignment controllability
political auditing
large language models
prompt steering
ideological dispersion
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt-based controllability
political alignment audit
steerability dispersion
ideological personas
LLM political bias
πŸ”Ž Similar Papers
No similar papers found.
B
Bartol Bućan
University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia
N
Nikola Sočec
University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia
S
Sarah Isufi
It From Bit d.o.o., Zagreb, Croatia
M
Morena Granić
University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia
L
Luka Hobor
University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia
A
Agneza Krajna
University of Zagreb Faculty of Electrical Engineering and Computing, Zagreb, Croatia
Mihael Kovac
Mihael Kovac
PhD Student, University of Zagreb, Faculty of electrical engineering and computing
geometric deep learningmathematical optimizationheterogeneous computing
Mario Brcic
Mario Brcic
Associate Professor, University of Zagreb Faculty of Electrical Engineering and Computing
Operations ResearchArtificial Intelligence