human-in-the-loop causal discovery

Designs and implements algorithms, interfaces, and workflows that elicit and incorporate human judgments to discover, refine, and validate causal graph structure. Uses active querying and interactive updates (e.g., updating Bayesian posteriors over structural causal models) to resolve ambiguous edges, reduce model uncertainty, estimate per-user causal effects, and improve interpretability and trust.

human-in-the-loopcausaldiscovery

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a Bayesian active learning framework for causal discovery under limited expert query budgets. The approach iteratively queries the existence and direction of local edges via triplet relationships to efficiently optimize the posterior over directed acyclic graphs (DAGs). It introduces a novel triplet likelihood model to account for noise in expert judgments and selects the most informative queries based on expected information gain. By integrating particle-based approximation with Bayesian inference, the method significantly improves DAG structure recovery accuracy on synthetic graphs, protein signaling networks, and human gene perturbation data. Moreover, it accelerates posterior convergence and enhances causal effect identification under stringent query constraints.

active queryingcausal discoverydirected acyclic graphs

In strongly self-selecting feedback systems, conventional causal discovery methods often fail to ensure reliable causal effect estimation due to their heavy reliance on the accuracy of recovered graph structures. This work proposes an effect-centric validation framework that prioritizes identifiability, treating causal graphs as testable structural hypotheses and evaluating them through identifiability, stability, and falsifiability. Moving beyond the prevailing paradigm that equates graph recovery accuracy with methodological validity, the framework emphasizes effect consistency tailored to specific causal queries and reveals a novel phenomenon: causal effect estimates can converge across distinct graph structures. Experiments on real-world game telemetry data demonstrate that algorithms satisfying identifiability conditions yield robust and consistent effect estimates, whereas those suffering from endpoint ambiguity produce unstable or attenuated effects.

causal discoverydecision supporteffect-level validation

Existing computational tools for qualitative data analysis often fall short in effectively supporting causal exploration due to insufficient contextual awareness, limited trustworthiness, or overly complex outputs. To address these limitations, this work proposes QualCausal, the first interactive causal analysis system grounded in user research–driven design principles. Developed through formative user studies, QualCausal integrates context-aware processing, cognitive scaffolding, and explainability mechanisms to facilitate efficient exploration and validation of causal hypotheses within qualitative datasets. The system enables researchers to extract causal relationships, construct interactive causal networks, and examine findings through coordinated multi-view visualizations. User evaluations demonstrate that QualCausal significantly reduces analytical burden, provides robust cognitive support, and prompts critical reflection on how computational tools can be meaningfully integrated into social science research practices, thereby bridging the gap between computational assistance and qualitative inquiry paradigms.

causal relationshipscomputational toolscontext

A Survey on Causal Discovery: Theory and Practice

May 17, 2023
AZ
Alessio Zanga
🏛️ University of Milano - Bicocca

This study addresses the challenge of identifying causal graphs and estimating causal effects from observational data. We propose the first unified analytical framework that horizontally integrates major causal discovery paradigms—including constraint-based methods (e.g., PC), score-based methods (e.g., GES), functional causal models (e.g., LiNGAM, ANM, CAM, NOTEARS), and neural causal learning—while rigorously characterizing their identifiability conditions and practical applicability boundaries. Our contribution comprises: (1) a comprehensive knowledge graph covering 12 algorithmic families, 8 open-source toolkits, and applications across six domains (e.g., healthcare, economics, ecology); (2) standardized benchmark datasets, reproducible evaluation protocols, and practitioner-oriented guidelines; and (3) paradigm-level unification, formal identification boundary analysis, and an end-to-end resource ecosystem for real-world causal discovery deployment.

Identifying causal effects using graphical modelsReviewing algorithms and tools for causal inferenceSurveying causal discovery methods from data

Existing causal discovery algorithms lack standardized evaluation benchmarks; conventional metrics (e.g., precision, recall) are highly susceptible to random guessing—especially in sparse graphs—yielding false positive rates exceeding 0.8 and severely compromising performance assessment. Method: We propose a normalized evaluation paradigm using random guessing as a negative control, focusing on skeleton estimation. We formally establish, for the first time, a theoretically grounded negative-control benchmark for causal discovery evaluation; derive exact distributions of key metrics under the random-hypothesis null; develop a statistically principled skeleton-fitting significance test grounded in statistical inference and random graph theory; and extend the negative-control framework to full causal structure evaluation. Contribution/Results: Validated via Monte Carlo simulations and real biological datasets, our approach substantially improves assessment reliability. We publicly release an open-source evaluation pipeline to foster community-wide standardization.

Lack of guidelines for evaluating causal discovery algorithmsNeed for negative controls to assess random guessing performanceProposing exact tests and pipelines for broader negative control use

Latest Papers

What's happening recently
View more

This work addresses the challenge that multiple agents in a shared system may hold divergent structural causal models (SCMs), leading to inconsistent judgments of fairness due to disagreements over interventional and counterfactual distributions. The paper introduces the first causally aware framework that distinguishes between structural and parametric causal awareness, proposes algorithms to compute interventional and counterfactual distributions, and employs probabilistic distance measures—such as KL divergence—to quantify discrepancies among agents. Experiments on the German Credit dataset demonstrate that causal awareness significantly influences both accuracy and fairness evaluations, with results highly sensitive to the choice of distance metric and decision threshold. These findings reveal the context-dependent nature of bias, challenge the assumption of a single objective fairness standard, and underscore the necessity of multi-perspective fairness analysis.

Causal PerceptionCounterfactual ReasoningFairness

本文针对现有因果推理框架无法处理循环因果依赖的问题,提出了一种新的二部图因果模型(BGCMs),通过明确指定干预方程、目标变量及其值来解决标准干预的模糊性。

causal reasoningcyclic causal dependenciesfeedback mechanisms

Existing recommendation-based remediation approaches often overlook individual users’ causal structures and feature interactions, making it difficult to generate personalized and actionable interventions. This work proposes a human-in-the-loop Bayesian causal inference framework that dynamically learns each user’s individual structural causal model through interactive queries and leverages this model to produce causally consistent, low-cost, and personalized remedial recommendations. By integrating human-in-the-loop mechanisms with Bayesian causal reasoning—a novel combination in this domain—the method generates more plausible and effective intervention strategies across both simulated linear and nonlinear causal environments.

algorithmic recoursecausal structurecounterfactual explanations

This work addresses the tendency of existing large language models to conflate textual associations or hallucinations with genuine causal evidence when applied to causal discovery, often lacking explicit grounding in data and assumptions. To remedy this, the authors propose a novel paradigm wherein AI agents assist—but do not autonomously generate—causal conclusions. This framework orchestrates data analysis, preprocessing, method recommendation, expert knowledge integration, and result interpretation to ensure that all causal inferences are rigorously based on empirical data, explicit assumptions, and formal algorithms. Built upon the causal-learn ecosystem, the team developed an online platform (causallearn.com) featuring agent-driven capabilities for data validation, context-aware retrieval, assumption articulation, and graphical model explanation. The efficacy and reliability of this human–AI collaborative approach are demonstrated through a case study on Big Five personality data.

agent-assisted analysiscausal discoverycausal evidence

Hot Scholars

MK

Matthias Kerzel

Knowledge Technology Engineer, University of Hamburg
AIArtificial Neural NetworksNeuroroboticsDevelopmental Robotics
RC

Ruichu Cai

Professor of Computer Science, Guangdong University of Technology
causality
KY

Kui Yu

Professor, Hefei University of Technology
Causal discovery and Data mining
AG

Adrian Groza

Technical University of Cluj-Napoca, European University of Technology (EUt+)
Artificial IntelligenceAgentic AIKnowledge representationExplainable AI