🤖 AI Summary
Large language models are prone to factual inaccuracies and ideological biases in political question-answering, with election-related contexts posing particularly high risks. This work proposes the first reproducible measurement framework that explicitly links hallucinated content to ideological orientation. Leveraging QBias—a novel dataset of expert-annotated political news—the study generates questions, detects reference-based contrastive hallucinations, and systematically evaluates model outputs using a fine-tuned stance classifier combined with logits-based uncertainty analysis. Findings reveal that while the ideological leaning of source material has limited impact on hallucination frequency, hallucinated content exhibits a significant leftward bias. Moreover, high-entropy generation contexts are more likely to induce hallucinations, and output uncertainty partially predicts this leftward drift, suggesting an “uncertainty-driven guessing” mechanism underlying ideological skew.
📝 Abstract
Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent information settings where factual errors and ideological distortions are high-stakes. We present a reproducible measurement framework that treats hallucinations, unsupported statements in document-grounded QA, as diagnostic signals of ideological drift. Using 21,727 expert-labeled U.S. political news articles from QBias spanning left, center, and right sources, we (i) generate an article-specific question, (ii) elicit document-grounded answers from three open-weight LLMs and one proprietary model, (iii) detect sentence-level hallucinations via reference-based comparison, (iv) classify the ideological valence of hallucinated sentences with a fine-tuned stance classifier, and (v) probe output logits to relate token-level uncertainty to hallucination and drift. Hallucination rates vary substantially across models and concentrate in contentious topics, while source-ideology differences in hallucination frequency are modest. In contrast, hallucination content exhibits robust leftward drift: a majority of hallucinated sentences are classified as left-leaning, including among hallucinations generated from right-leaning sources. Logit-level analysis shows hallucinations arise in high-entropy generation contexts, and in some models uncertainty also predicts leftward drift, consistent with an"uncertainty to guessing"mechanism. We discuss implications for auditing AI-mediated political information and for designing safeguards in election-relevant deployments.