Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the pervasive hallucination problem in large chemical reasoning models, where structural descriptions in chain-of-thought (CoT) reasoning often diverge from actual molecular structures and are decoupled from answer correctness. Through systematic analysis of chemical CoT mechanisms, the study reveals for the first time that CoT exhibits both hallucinatory and functional characteristics, challenging the conventional view that treats CoT as reliable evidence of valid reasoning. Employing attribution analysis, SMILES-based draft perturbations, and cross-model comparisons across four model families and twelve chemical tasks, the authors demonstrate that correct answers frequently co-occur with inaccurate structural descriptions, while perturbing the molecular drafts significantly degrades output quality. These findings indicate that CoT plays a causal role in generation, underscoring the critical importance of process supervision in chemical reasoning.
📝 Abstract
Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.
Problem

Research questions and friction points this paper is trying to address.

hallucination
chain-of-thought
chemical reasoning
molecular structure
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought
Hallucination
Molecular Scratchpad
SMILES
Attribution Analysis
🔎 Similar Papers