Mapping and Advancing the Scalability-Accuracy Frontier of Nonlinear Causal Discovery

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the scalability-accuracy trade-off bottleneck in nonlinear causal discovery by systematically comparing four mainstream paradigms: differentiable structure learning, amortized inference, score matching, and combinatorial search. We propose SPADE, a novel spline-based scoring scheme that leverages precompiled statistic reuse to reduce the computational complexity of Gaussian variants from O(ndΒ³) to O(ndΒ² + dΒ³). Experimental results demonstrate that this approach achieves solution times on the order of seconds to minutes for graphs comprising thousands of nodes. By substantially expanding the scalability-accuracy Pareto frontier of causal discovery, this work offers a computationally efficient and highly accurate framework for large-scale nonlinear causal structure learning.
πŸ“ Abstract
Scalable nonlinear causal discovery requires methods that combine flexible mechanism estimators with efficient search over large graph spaces. Several algorithmic families have been proposed to address this challenge, yet their accuracy-runtime trade-offs remain poorly understood. We empirically compare the four major approaches: differentiable structure learning, amortized structure learning, score-matching, and combinatorial search. Our results reveal complementary bottlenecks: differentiable and amortized methods scale well but exhibit an accuracy gap, score-matching methods can be accurate in low dimensions but degrade quickly for increasing feature sizes, and combinatorial methods remain accurate but are slowed by repeated and redundant local scoring. Motivated by this bottleneck, we develop SPADE, a spline-based score-evaluation scheme that compiles sufficient statistics once and reuses them throughout combinatorial search. Under bounded indegree, its Gaussian variant reduces algorithmic complexity from O(nd^3) to O(nd^2+d^3). Empirically, SPADE shifts the observed scalability-accuracy frontier by orders of magnitude: it solves 100-variable problems with 160K samples in seconds and 1600-variable problems with 2.5K samples in minutes, while retaining high structural accuracy across synthetic and real-world benchmarks. These results reveal a substantial shift in the practical scale of combinatorial search and highlight the importance of evaluating scalable causal-discovery methods along the full accuracy-runtime frontier.
Problem

Research questions and friction points this paper is trying to address.

nonlinear causal discovery
scalability-accuracy trade-off
combinatorial search
structure learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nonlinear Causal Discovery
Combinatorial Search
Spline-based Score Evaluation
Scalability-Accuracy Frontier
Sufficient Statistics Compilation
πŸ’Ό Related Jobs
No related jobs found.
H
Hendrik Suhr
CISPA Helmholtz Center for Information Security
S
Sascha Xu
CISPA Helmholtz Center for Information Security
Jilles Vreeken
Jilles Vreeken
CISPA Helmholtz Center for Information Security
Machine LearningCausal InferenceData Mining