π€ AI Summary
This study addresses a critical safety vulnerability in large language models (LLMs), demonstrating that their decision-making is systematically biased when passively exposed to external, task-irrelevant content. To investigate this phenomenon, the work presents the first systematic quantification of the irrational influence exerted by passive exposure on model decisions. Through cross-model comparisons, real-world web opinion injection, and logical consistency evaluations, it examines how irrelevant contexts disrupt model judgment and induce instruction violations. The findings reveal a novel risk dimension wherein models are highly susceptible to manipulation by meaningless contexts. Notably, experiments show that closed-weight models exhibit a decision bias rate approaching 50%, confirming that passive exposure can lead LLMs to accept false claims and violate explicit user instructions.
π Abstract
Large language model (LLM) assistants can now search the web and consult external sources while completing user requests. These sources can provide useful evidence, but they can also introduce additional content into the model's context. Can such passive exposure steer a decision even when the added content provides no reason to change it? We examine the stability of model decisions on the same tasks with and without such external content. Across all open-weight and closed-weight models we test, exposure systematically shifts decisions, with effects reaching nearly 50 percentage points in closed-weight models. The same pattern appears with real-world online opinions. The influence also extends beyond subjective preferences. Such exposure can steer models toward choices that violate explicit user requirements and increase their acceptance of false claims. In short, what enters an LLM's context can influence its decision even when it should not determine it.