llm forensic reasoning

Designs and evaluates systems that use large language models to reconstruct, explain, and prioritize causes and intents behind observed events or artifacts—translating raw alerts or signals into high-level intents, generating causal hypotheses, and retrieving and citing supporting evidence. Builds workflows for adversarial deliberation and counterfactual reasoning so hypotheses can be tested and refined and so outputs include traceable chains of reasoning.

llmforensicreasoning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.26
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

On the Eligibility of LLMs for Counterfactual Reasoning: A Decompositional Study

May 17, 2025
SY
Shuai Yang
🏛️ Binghamton University | Shanghai University

Large language models (LLMs) exhibit inconsistent performance in counterfactual reasoning, and the underlying causes—particularly the role of modality dependence—remain poorly understood. Method: We propose the first decomposable evaluation framework for counterfactual reasoning, disentangling the task into two sequential stages: *causal structure construction* and *counterfactual intervention reasoning*. We conduct systematic benchmarking across 11 cross-modal datasets spanning text, mathematics, code, and vision-language domains. Contribution/Results: Through stage-wise behavioral analysis, we uncover— for the first time—the critical influence of modality type and intermediate reasoning steps on LLMs’ counterfactual capabilities. We precisely identify *causal modeling* (not intervention reasoning) as the primary bottleneck. Our framework establishes an interpretable diagnostic pathway and provides both theoretical grounding and concrete optimization directions for enhancing LLMs’ robust counterfactual reasoning.

Decompose counterfactual reasoning into causality and intervention stagesEvaluate LLMs across 11 diverse tasks and modalitiesIdentify factors impeding LLMs' counterfactual reasoning performance

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?

Jun 26, 2025
HC
Haoang Chi
🏛️ National University of Defense Technology | University of Melbourne | Intelligent Game and Decision Lab | University of Sydney | Hong Kong Baptist University

This study investigates whether large language models (LLMs) possess human-like level-2 causal reasoning—grounded in structural causal models and counterfactual reasoning—beyond superficial, correlation-based level-1 inference. Method: We propose G²-Reasoner, a framework integrating general-knowledge injection with goal-directed prompting to guide structured causal modeling. Leveraging the autoregressive nature of Transformer architectures, we introduce CausalProbe-2024, the first causal QA benchmark explicitly designed to evaluate novelty-aware and counterfactual reasoning. Contribution/Results: Experiments demonstrate that G²-Reasoner significantly outperforms baselines across cross-domain causal identification and counterfactual intervention inference. It validates a viable pathway from parameterized, experience-driven inference toward interpretable, generalizable level-2 causal reasoning. This work establishes a novel paradigm for advancing LLMs toward genuine causal intelligence.

Assessing if LLMs perform genuine human-like causal reasoningIdentifying LLMs' limitation to shallow level-1 causal reasoningProposing a method to enhance LLMs' causal reasoning capabilities

Leveraging Large Language Models for Automated Causal Loop Diagram Generation: Enhancing System Dynamics Modeling through Curated Prompting Techniques

Mar 23, 2025
NG
Ning-Yuan Georgia Liu
🏛️ Harvard Medical School | University of Melbourne | Massachusetts Institute of Technology

Causal Loop Diagram (CLD) construction in system dynamics suffers from low efficiency and high entry barriers for novices. Method: This paper proposes the first stepwise prompt engineering framework tailored for CLD generation, leveraging large language models (LLMs) to automatically map textual dynamic hypotheses into structured CLDs. The approach integrates chain-of-thought reasoning, role-guided prompting, and domain-specific constraints, representing CLDs as standard directed graphs; it is fine-tuned and evaluated on a textbook-based system dynamics dataset. Contribution/Results: Experiments show that the automatically generated CLDs achieve 89% agreement with expert-built diagrams on simple dynamic structures, substantially reducing modeling time. This work establishes the first end-to-end, accurate, interpretable, and domain-aligned natural-language-to-CLD generation pipeline, empirically validating the feasibility and practical utility of LLMs in automating system modeling.

Automating causal loop diagram generation from dynamic hypothesesEvaluating LLM performance in creating expert-quality CLDs with curated promptsOvercoming challenges in extracting variables and relationships for novice modelers

Reasoning Elicitation in Language Models via Counterfactual Feedback

Oct 02, 2024
AH
Alihan Hüyük
🏛️ Harvard University | Microsoft Research | Cornell Tech

Large language models exhibit limited capability in causal reasoning tasks—particularly counterfactual question answering—due to inherent biases and insufficient grounding in causal mechanisms. Method: We propose a novel paradigm for enhancing causal reasoning: (1) introducing CausalQA-Balanced, the first evaluation metric jointly optimizing factual and counterfactual accuracy to quantify reasoning bias; and (2) designing a causal-mechanism-inspired fine-tuning strategy integrating counterfactual question generation, multi-objective supervised fine-tuning, and feedback-driven optimization. Contribution/Results: Our approach significantly improves model accuracy on counterfactual QA and strengthens generalization across inductive, deductive, and cross-task causal reasoning. Extensive experiments demonstrate systematic superiority over state-of-the-art baselines across multiple real-world scenarios, establishing a new benchmark for causally grounded language understanding.

Develop metrics for factual and counterfactual reasoning accuracy.Enhance reasoning in language models via counterfactual feedback.Improve generalization in inductive and deductive reasoning tasks.

CausalEval: Towards Better Causal Reasoning in Language Models

Oct 22, 2024
SX
Siheng Xiong
🏛️ Georgia Institute of Technology | University of Massachusetts Amherst | University of California, Los Angeles | University of California, San Diego | University of Arizona | Arizona State University | Michigan State University | Purdue University

Large language models (LLMs) exhibit weak performance on core causal reasoning tasks—such as counterfactual inference and intervention analysis—with accuracy consistently below 65%. To address this, we systematically evaluate LLMs’ causal reasoning capabilities and propose a dual-path taxonomy: (1) LLMs as reasoning engines, and (2) LLMs as knowledge- or data-augmented assistants. We introduce CausalEval, the first unified benchmark for causal reasoning evaluation. Methodologically, CausalEval integrates diverse tasks—including causal discovery, do-calculus, and counterfactual reasoning—and employs structured causal prompting, causal-graph-guided fine-tuning, and neuro-symbolic hybrid modeling. Empirical results demonstrate that knowledge injection techniques yield substantial improvements, with structured causal prompting and causal-graph-guided fine-tuning emerging as the most effective approaches. This work establishes a reproducible benchmark, a principled classification framework, and empirically grounded insights for assessing and enhancing causal reasoning in LLMs.

Enhancing language models' causal reasoning capabilitiesEvaluating current models on causal reasoning tasksIdentifying future research directions in causal reasoning

Latest Papers

What's happening recently
View more

This work proposes a theory-guided hypothesis generation framework that integrates scientific explanatory mechanisms with large language models (LLMs) to systematically derive novel, testable hypotheses from scientific literature. Addressing the limitations of traditional approaches—which struggle to efficiently and rigorously generate hypotheses from existing research—the method begins with published conclusions, reconstructs their underlying theoretical foundations, and produces structured new hypotheses. By synergizing LLMs, scientific explanation reconstruction, LLM-as-judge evaluation, and expert validation, the framework significantly outperforms direct hypothesis-generation baselines in data science. Notably, two high-scoring generated hypotheses were implemented as novel algorithms, both surpassing the original baseline models in performance and demonstrating strong cross-domain generalization potential.

AI-driven workflowhypothesis generationlarge language models

Large language models exhibit fragility in counterfactual reasoning, reflecting a lack of robust causal inference capabilities. To address this limitation, this work proposes Dual Counterfactual Consistency (DCC), a novel inference-time mechanism that enables causal evaluation and enhancement without requiring additional training or annotated data. DCC constructs dual counterfactual scenarios and integrates test-time rejection sampling to guide the model in performing causal interventions and counterfactual predictions. Extensive experiments across multiple mainstream large language models and diverse causal reasoning benchmarks demonstrate that DCC significantly improves causal reasoning performance, thereby validating its effectiveness and generalizability.

causal reasoningcounterfactual reasoninglarge language models

This study evaluates the hypothesis-driven inductive reasoning capabilities of large language models (LLMs) in the context of scientific discovery. Inspired by the Wason 2-4-6 task, we introduce an interactive rule-discovery benchmark that requires models to propose examples, receive feedback, and iteratively uncover hidden rules—thereby simulating the core scientific reasoning processes of hypothesis generation, evidence gathering, and belief revision. For the first time, hypothesis testing and falsification behaviors are formally incorporated into the quantitative assessment of LLMs’ scientific reasoning. Experiments across twelve mainstream models reveal that those with explicit reasoning mechanisms outperform purely instruction-tuned counterparts, and models actively engaging in falsification tests achieve significantly better performance. Nevertheless, overall results remain substantially below ideal levels, exposing fine-grained failure modes in hypothesis-space exploration.

falsificationhypothesis testinginductive reasoning

This work proposes the first systematic framework for treating large language models (LLMs) as implicit sources of causal knowledge, aiming to extract their latent assumptions about causal relationships among events within a given topic. The approach involves generating topic-relevant text, extracting and normalizing events, constructing binary event indicator vectors, and applying causal discovery algorithms to infer underlying causal graph structures. By representing LLMs’ internal causal beliefs in terms of testable variables and graphical models, the framework offers a novel and interpretable means of externalizing these implicit assumptions. Empirical results demonstrate that the method successfully produces a set of plausible, interpretable candidate causal graphs, thereby validating the presence of coherent causal reasoning embedded within the model’s representations.

causal discoverycausal graphscausality

Hot Scholars

FL

Feifei Li

Alibaba Group/Alibaba Cloud
DatabasesDatabaseData AnalyticsSystems
MS

Mo Sha

Associate Professor, Florida International University
Wireless NetworkingInternet of ThingsApplied Machine LearningNetwork Security
TX

Tianyin Xu

University of Illinois at Urbana-Champaign
Software/system reliabilityOperating systemsDistributed systemsSoftware engineering
FY

Fan Yang

Microsoft Research
Systems
PB

Paul Buitelaar

Professor in Data Analytics, Data Science Institute, Univ of Galway, Co-PI Insight Centre
Natural Language ProcessingKnowledge GraphsText MiningSemantics