Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study identifies a systemic hallucination risk in large language models (LLMs) deployed for news editing—termed *explanatory overconfidence*: models frequently generate authoritative, definitive assertions lacking documentary support, violating journalism’s core tenets of source attribution and claim verifiability. We develop a news-practice-oriented hallucination taxonomy and evaluate ChatGPT, Gemini, and NotebookLM on 300 U.S. TikTok litigation and policy documents, systematically varying prompt specificity and context window size. Results show an overall hallucination rate of 30%, with ChatGPT and Gemini exhibiting markedly higher rates (~40%) than NotebookLM (13%); errors predominantly involve unsupported generalizations rather than entity fabrication. The paper introduces the novel concept of *explanatory overconfidence* and advocates architectural redesign prioritizing mandatory source attribution over textual fluency. Findings provide both theoretical grounding and empirical evidence for designing trustworthy AI tools in journalistic practice.

Technology Category

Natural Language Processing: Fact-Checking / Misinformation Detection (NLP Focus)Machine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Large language models (LLMs) are increasingly used in newsroom workflows, but their tendency to hallucinate poses risks to core journalistic practices of sourcing, attribution, and accuracy. We evaluate three widely used tools - ChatGPT, Gemini, and NotebookLM - on a reporting-style task grounded in a 300-document corpus related to TikTok litigation and policy in the U.S. We vary prompt specificity and context size and annotate sentence-level outputs using a taxonomy to measure hallucination type and severity. Across our sample, 30% of model outputs contained at least one hallucination, with rates approximately three times higher for Gemini and ChatGPT (40%) than for NotebookLM (13%). Qualitatively, most errors did not involve invented entities or numbers; instead, we observed interpretive overconfidence - models added unsupported characterizations of sources and transformed attributed opinions into general statements. These patterns reveal a fundamental epistemological mismatch: While journalism requires explicit sourcing for every claim, LLMs generate authoritative-sounding text regardless of evidentiary support. We propose journalism-specific extensions to existing hallucination taxonomies and argue that effective newsroom tools need architectures that enforce accurate attribution rather than optimize for fluency.
Problem

Research questions and friction points this paper is trying to address.

Evaluating LLM hallucination risks in journalistic document-based queries
Measuring interpretive overconfidence in source characterization and attribution
Addressing epistemological mismatch between journalism standards and LLM outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluated three LLM tools on document-based queries
Proposed journalism-specific hallucination taxonomy extensions
Recommended architectures enforcing attribution over fluency
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
N
Nick Hagar
Northwestern University, Evanston, IL, USA
W
Wilma Agustianto
University of Minnesota, Minneapolis, MN, USA
Nicholas Diakopoulos
Nicholas Diakopoulos
Professor, Northwestern University
Computational JournalismAlgorithmic AccountabilityHuman Computer InteractionAI EthicsSocial