Score
Designs and evaluates systems that use large language models to reconstruct, explain, and prioritize causes and intents behind observed events or artifacts—translating raw alerts or signals into high-level intents, generating causal hypotheses, and retrieving and citing supporting evidence. Builds workflows for adversarial deliberation and counterfactual reasoning so hypotheses can be tested and refined and so outputs include traceable chains of reasoning.
Large language models (LLMs) exhibit inconsistent performance in counterfactual reasoning, and the underlying causes—particularly the role of modality dependence—remain poorly understood. Method: We propose the first decomposable evaluation framework for counterfactual reasoning, disentangling the task into two sequential stages: *causal structure construction* and *counterfactual intervention reasoning*. We conduct systematic benchmarking across 11 cross-modal datasets spanning text, mathematics, code, and vision-language domains. Contribution/Results: Through stage-wise behavioral analysis, we uncover— for the first time—the critical influence of modality type and intermediate reasoning steps on LLMs’ counterfactual capabilities. We precisely identify *causal modeling* (not intervention reasoning) as the primary bottleneck. Our framework establishes an interpretable diagnostic pathway and provides both theoretical grounding and concrete optimization directions for enhancing LLMs’ robust counterfactual reasoning.
This study investigates whether large language models (LLMs) possess human-like level-2 causal reasoning—grounded in structural causal models and counterfactual reasoning—beyond superficial, correlation-based level-1 inference. Method: We propose G²-Reasoner, a framework integrating general-knowledge injection with goal-directed prompting to guide structured causal modeling. Leveraging the autoregressive nature of Transformer architectures, we introduce CausalProbe-2024, the first causal QA benchmark explicitly designed to evaluate novelty-aware and counterfactual reasoning. Contribution/Results: Experiments demonstrate that G²-Reasoner significantly outperforms baselines across cross-domain causal identification and counterfactual intervention inference. It validates a viable pathway from parameterized, experience-driven inference toward interpretable, generalizable level-2 causal reasoning. This work establishes a novel paradigm for advancing LLMs toward genuine causal intelligence.
Causal Loop Diagram (CLD) construction in system dynamics suffers from low efficiency and high entry barriers for novices. Method: This paper proposes the first stepwise prompt engineering framework tailored for CLD generation, leveraging large language models (LLMs) to automatically map textual dynamic hypotheses into structured CLDs. The approach integrates chain-of-thought reasoning, role-guided prompting, and domain-specific constraints, representing CLDs as standard directed graphs; it is fine-tuned and evaluated on a textbook-based system dynamics dataset. Contribution/Results: Experiments show that the automatically generated CLDs achieve 89% agreement with expert-built diagrams on simple dynamic structures, substantially reducing modeling time. This work establishes the first end-to-end, accurate, interpretable, and domain-aligned natural-language-to-CLD generation pipeline, empirically validating the feasibility and practical utility of LLMs in automating system modeling.
Large language models exhibit limited capability in causal reasoning tasks—particularly counterfactual question answering—due to inherent biases and insufficient grounding in causal mechanisms. Method: We propose a novel paradigm for enhancing causal reasoning: (1) introducing CausalQA-Balanced, the first evaluation metric jointly optimizing factual and counterfactual accuracy to quantify reasoning bias; and (2) designing a causal-mechanism-inspired fine-tuning strategy integrating counterfactual question generation, multi-objective supervised fine-tuning, and feedback-driven optimization. Contribution/Results: Our approach significantly improves model accuracy on counterfactual QA and strengthens generalization across inductive, deductive, and cross-task causal reasoning. Extensive experiments demonstrate systematic superiority over state-of-the-art baselines across multiple real-world scenarios, establishing a new benchmark for causally grounded language understanding.
Large language models (LLMs) exhibit weak performance on core causal reasoning tasks—such as counterfactual inference and intervention analysis—with accuracy consistently below 65%. To address this, we systematically evaluate LLMs’ causal reasoning capabilities and propose a dual-path taxonomy: (1) LLMs as reasoning engines, and (2) LLMs as knowledge- or data-augmented assistants. We introduce CausalEval, the first unified benchmark for causal reasoning evaluation. Methodologically, CausalEval integrates diverse tasks—including causal discovery, do-calculus, and counterfactual reasoning—and employs structured causal prompting, causal-graph-guided fine-tuning, and neuro-symbolic hybrid modeling. Empirical results demonstrate that knowledge injection techniques yield substantial improvements, with structured causal prompting and causal-graph-guided fine-tuning emerging as the most effective approaches. This work establishes a reproducible benchmark, a principled classification framework, and empirically grounded insights for assessing and enhancing causal reasoning in LLMs.
This work proposes a theory-guided hypothesis generation framework that integrates scientific explanatory mechanisms with large language models (LLMs) to systematically derive novel, testable hypotheses from scientific literature. Addressing the limitations of traditional approaches—which struggle to efficiently and rigorously generate hypotheses from existing research—the method begins with published conclusions, reconstructs their underlying theoretical foundations, and produces structured new hypotheses. By synergizing LLMs, scientific explanation reconstruction, LLM-as-judge evaluation, and expert validation, the framework significantly outperforms direct hypothesis-generation baselines in data science. Notably, two high-scoring generated hypotheses were implemented as novel algorithms, both surpassing the original baseline models in performance and demonstrating strong cross-domain generalization potential.
Large language models exhibit fragility in counterfactual reasoning, reflecting a lack of robust causal inference capabilities. To address this limitation, this work proposes Dual Counterfactual Consistency (DCC), a novel inference-time mechanism that enables causal evaluation and enhancement without requiring additional training or annotated data. DCC constructs dual counterfactual scenarios and integrates test-time rejection sampling to guide the model in performing causal interventions and counterfactual predictions. Extensive experiments across multiple mainstream large language models and diverse causal reasoning benchmarks demonstrate that DCC significantly improves causal reasoning performance, thereby validating its effectiveness and generalizability.
This study evaluates the hypothesis-driven inductive reasoning capabilities of large language models (LLMs) in the context of scientific discovery. Inspired by the Wason 2-4-6 task, we introduce an interactive rule-discovery benchmark that requires models to propose examples, receive feedback, and iteratively uncover hidden rules—thereby simulating the core scientific reasoning processes of hypothesis generation, evidence gathering, and belief revision. For the first time, hypothesis testing and falsification behaviors are formally incorporated into the quantitative assessment of LLMs’ scientific reasoning. Experiments across twelve mainstream models reveal that those with explicit reasoning mechanisms outperform purely instruction-tuned counterparts, and models actively engaging in falsification tests achieve significantly better performance. Nevertheless, overall results remain substantially below ideal levels, exposing fine-grained failure modes in hypothesis-space exploration.
This work proposes the first systematic framework for treating large language models (LLMs) as implicit sources of causal knowledge, aiming to extract their latent assumptions about causal relationships among events within a given topic. The approach involves generating topic-relevant text, extracting and normalizing events, constructing binary event indicator vectors, and applying causal discovery algorithms to infer underlying causal graph structures. By representing LLMs’ internal causal beliefs in terms of testable variables and graphical models, the framework offers a novel and interpretable means of externalizing these implicit assumptions. Empirical results demonstrate that the method successfully produces a set of plausible, interpretable candidate causal graphs, thereby validating the presence of coherent causal reasoning embedded within the model’s representations.