🤖 AI Summary
This study addresses the "audit gap" in language models induced by triggerless misinformation, wherein factual responses remain correct yet downstream decisions become erroneous. Methodologically, the authors construct controlled decision tasks (Guess the Capital) via data poisoning and analyze a real-world Facebook case to systematically evaluate the differential impacts of poisoned training data on factual question answering versus subsequent decision-making across multiple models. The core contribution lies in the first quantification of this audit gap: experiments reveal that while probing-based detection remains effective under high injection rates, significant decision biases persist. Furthermore, even after correcting the model's factual outputs, erroneous decisions continue to occur, demonstrating that relying solely on fact-checking is insufficient to ensure decision safety.
📝 Abstract
Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a misleading claim from Facebook posts about the 2019--20 Australian bushfires. Across eight models at dose 1,000, direct injected-choice rates reach 95.8--100\%, while injected game choices increase by 1.7--14.4 percentage points over matched truthful training. The gap runs the other way too. Facts that pass the direct probe still push decisions toward the injected answer, and game accuracy drops further than those choices explain. Truthful correction brings the fact back but not the decisions built on it. We then look into the real-world bushfire case, models trained on the false posts say that people were arrested for arson even when they lose the inflated count, and in a count-by-wording factorial the misleading arrest wording produces arrest assertions even when the training count stays at 24. In simpler terms, \textbf{a correct factual answer does not guarantee a correct decision, and losing the injected number does not remove the misleading story}.