🤖 AI Summary
This work addresses the longstanding reliance on manual curation in constructing auditable, literature-grounded causal directed acyclic graphs (DAGs) for biomedical causal analysis, which has lacked systematic computational support. The authors propose the first browser-based system that integrates large language models with verifiable literature evidence to automatically generate structured causal judgments from free-text input. The system retrieves document snapshots, extracts causal evidence, and produces fully traceable DAGs that satisfy constraint consistency. Evaluated on a literature-based benchmark, the approach significantly outperforms pure large language model baselines, achieving high edge recall while preserving a complete, inspectable chain of evidence to ensure expert auditability throughout the causal modeling process.
📝 Abstract
Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. Analysts must connect study variables to prior literature, evaluate uncertain causal claims, and preserve sufficient provenance for expert review. We present DAGForge, a browser-based system for authoring causal DAGs as auditable, evidence-linked artifacts. Given free-text descriptions of study concepts, DAGForge creates a reproducible literature snapshot, uses an LLM-based reasoning module to generate structured pairwise causal judgments grounded in verbatim evidence excerpts, and assembles those judgments into a constraint-checked graph. Each proposed edge includes confidence estimates, provenance, and a reviewable rationale. The interface supports study specification, progress monitoring, evidence review, graph comparison, adjustment-set computation, and export. In evaluations against both compact benchmark DAGs and reference DAGs derived from published literature, DAGForge achieves high edge recall on the literature-based cohort while retaining verifiable evidence trails absent from LLM-only baselines. DAGForge thus reduces the burden of causal DAG curation while making the resulting assumptions auditable, supporting the design, analysis, and interpretation of biomedical studies.