Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features

📅 2025-05-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper identifies the root cause of frequent failures by large language models (LLMs) on constraint-satisfaction reasoning tasks—such as graph coloring—not as general factual inaccuracies, but as systematic hallucination of unspecified edges in input graph structures. Method: Through cross-model error analysis (o1-mini, DeepSeek-R1, Claude 3.7), chain-of-thought tracing, and controlled experiments varying variable complexity and semantic framing, we isolate and quantify structural hallucination. Contribution/Results: We provide the first systematic empirical evidence that “structural feature hallucination” is the dominant failure mechanism in logical reasoning, with over 80% of errors attributable to hallucinated edges in certain models. We further propose a novel structure-aware training and verification paradigm, grounded in formal graph semantics, to mitigate such hallucinations. This work establishes both a theoretical foundation and practical methodology for enhancing LLM robustness in structured reasoning tasks.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Computational Complexity of ReasoningNatural Language Processing: (Large) Language Models

Application Category

Graph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for search
📝 Abstract
Large language models have recently made great strides in reasoning task performance through chain-of-thought (CoT) strategies trained via reinforcement learning; however, these"reasoning large language models"(RLLMs) remain imperfect reasoners, and understanding the frequencies and causes of their failure modes is important for both users and developers. We test o1-mini, o3-mini, DeepSeek-R1, Claude 3.7 Sonnet, Gemini 2.5 Pro Preview, and Grok 3 Mini Beta on graph coloring as a variable-complexity constraint-satisfaction logic problem, and find evidence from both error rate comparisons and CoT/explanation text analysis that RLLMs are prone to hallucinate edges not specified in the prompt's description of the graph. This phenomenon persists across multiple problem complexity levels and semantic frames, and it appears to account for a significant fraction of the incorrect answers from every tested model, and the vast majority of them for some models. Our results indicate that RLLMs may possess broader issues with misrepresentation of problem specifics, and we offer suggestions for design choices to mitigate this weakness.
Problem

Research questions and friction points this paper is trying to address.

RLLMs hallucinate unspecified graph edges in reasoning tasks
Error rates show persistent misrepresentation of problem specifics
Design suggestions aim to mitigate hallucination-related failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using chain-of-thought strategies via reinforcement learning
Testing models on graph coloring constraint-satisfaction problems
Analyzing hallucinated edges in model explanations
🔎 Similar Papers
No similar papers found.