🤖 AI Summary
This paper identifies the root cause of frequent failures by large language models (LLMs) on constraint-satisfaction reasoning tasks—such as graph coloring—not as general factual inaccuracies, but as systematic hallucination of unspecified edges in input graph structures. Method: Through cross-model error analysis (o1-mini, DeepSeek-R1, Claude 3.7), chain-of-thought tracing, and controlled experiments varying variable complexity and semantic framing, we isolate and quantify structural hallucination. Contribution/Results: We provide the first systematic empirical evidence that “structural feature hallucination” is the dominant failure mechanism in logical reasoning, with over 80% of errors attributable to hallucinated edges in certain models. We further propose a novel structure-aware training and verification paradigm, grounded in formal graph semantics, to mitigate such hallucinations. This work establishes both a theoretical foundation and practical methodology for enhancing LLM robustness in structured reasoning tasks.
📝 Abstract
Large language models have recently made great strides in reasoning task performance through chain-of-thought (CoT) strategies trained via reinforcement learning; however, these"reasoning large language models"(RLLMs) remain imperfect reasoners, and understanding the frequencies and causes of their failure modes is important for both users and developers. We test o1-mini, o3-mini, DeepSeek-R1, Claude 3.7 Sonnet, Gemini 2.5 Pro Preview, and Grok 3 Mini Beta on graph coloring as a variable-complexity constraint-satisfaction logic problem, and find evidence from both error rate comparisons and CoT/explanation text analysis that RLLMs are prone to hallucinate edges not specified in the prompt's description of the graph. This phenomenon persists across multiple problem complexity levels and semantic frames, and it appears to account for a significant fraction of the incorrect answers from every tested model, and the vast majority of them for some models. Our results indicate that RLLMs may possess broader issues with misrepresentation of problem specifics, and we offer suggestions for design choices to mitigate this weakness.