🤖 AI Summary
This study addresses the issue that free-text chain-of-thought reasoning in large language models (LLMs) often leads to matching errors among Markov equivalence classes during causal inference. To overcome the limitations of unstructured intermediate reasoning, we propose a structured two-stage framework adhering to the principle of externalizing latent objects and constraining their forms. The method first guides the model to generate typed CPDAG summaries, then answers causal queries based on these graph states. By integrating the PC algorithm with schema constraint techniques, experiments on Qwen and GPT series models demonstrate a significant 13.4 percentage point improvement in F1 score (p<0.05), along with strong out-of-distribution robustness. These results validate the effectiveness of externalizing graph structures for enhancing the causal reasoning capabilities of LLMs.
📝 Abstract
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause full test, Structured Thinking raises Qwen3.5-27B from $73.0$ to $86.4$ $F_1$(Yes) over a strong PC-instruction baseline in the primary paired run ($+13.4$ pp; McNemar $p=2.4\times 10^{-6}$; bootstrap $95\%$ CI [$+8.4$, $+18.6$]); across three full-ID seeds, the mean gain is $+8.1 \pm 5.3$ pp. A PC-scaffolded two-turn prose control reaches only $67.6$ $F_1$, indicating that a detailed PC scaffold plus a schema-free prose intermediate is not sufficient. The same pattern holds on Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini. Scrambling the emitted CPDAG costs $12.0$ pp $F_1$, and a full-split audit shows close agreement with the reference CPDAG (ID skeleton $F_1$ $0.960$; exact match $75.9\%$). These results support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers use it.