Nonnegative DAG Learning via Concomitant Estimation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of learning directed acyclic graphs (DAGs) with non-negative edge weights from observational data. To this end, it proposes the NoCo estimator, which jointly recovers the weighted graph structure and noise variances. By exploiting the non-negativity of edge weights, the method simplifies the optimization landscape. Furthermore, it introduces a smoothed companion Lasso criterion coupled with a log-determinant constraint to circumvent KKT degeneracy. An efficient solver is developed based on the method of multipliers, block sequential convex approximation, and proximal gradient descent. The proposed approach significantly improves both DAG structure recovery and edge weight estimation accuracy across diverse scenarios, offering a novel framework for non-negative causal discovery that combines rigorous theoretical guarantees with computational efficiency.
📝 Abstract
We study the problem of learning directed acyclic graphs (DAGs) with nonnegative edge weights from observational data. We propose the Nonnegative and Concomitant (NoCo) DAG estimator, which jointly recovers the weighted graph structure and the exogenous noise variances in the linear structural equation model for the observations. Different from prior art, this noise-adaptive formulation blends a smoothed concomitant lasso criterion with a simpler log-determinant acyclicity constraint that exploits nonnegativity and yields a more benign optimization landscape. Specifically, nonnegative weights allow us to impose acyclicity directly on the adjacency matrix without elementwise squaring of its entries, thus avoiding the well-documented degeneracy of the Karush-Kuhn-Tucker conditions. Computationally, we develop a method of multipliers' algorithm that accommodates both homoscedastic and heteroscedastic noise profiles. Within each iteration, we use block successive convex approximation to minimize the augmented Lagrangian, alternating between proximal gradient steps for the adjacency matrix and closed-form updates for the noise scales. Simulated experiments demonstrate NoCo's improved structural and edge-weight recovery relative to competing methods in a variety of settings, highlighting the benefits of exploiting nonnegativity along with noise adaptivity.
Problem

Research questions and friction points this paper is trying to address.

Directed Acyclic Graph
Nonnegative Edge Weights
Observational Data
Structure Learning
Linear Structural Equation Model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Directed Acyclic Graph
Concomitant Estimation
Nonnegative Weights
Log-determinant Acyclicity Constraint
Block Successive Convex Approximation