🤖 AI Summary
This study addresses the non-identifiability inherent in multi-stream distribution models, where multiple streams act on a common distribution, leading to flat likelihood surfaces and unstable estimation. To resolve this, the authors introduce structured regularization into the framework for the first time, proposing a penalized maximum likelihood estimator with stream-specific heterogeneous penalties that integrate non-smooth regularizers such as adaptive Lasso and elastic net. An efficient proximal gradient algorithm is developed for optimization. Theoretical analysis demonstrates that the proposed method restores model identifiability, ensures strict convexity of the objective function for a unique solution, and establishes oracle properties for the adaptive Lasso component. Simulations and an empirical analysis of NHANES data on asthma and lead exposure show that the method substantially reduces estimation error, controls false discovery rates, and yields stable, interpretable estimates of protective effects.
📝 Abstract
Regression by composition provides a flexible framework for constructing conditional distributions through sequential group actions. However, when multiple flows act on the same distribution, the model becomes non-identifiable, leading to flat likelihood regions and unstable estimates. We introduce a structured regularization framework that resolves this issue by assigning flow-specific penalties. The resulting estimator is defined as a penalized maximum likelihood problem with heterogeneous regularization across flows. We establish theoretical properties, including identifiability under penalization, uniqueness of the minimizer via strict convexification, and asymptotic consistency. For the adaptive Lasso, we further prove the oracle property. An efficient proximal gradient algorithm handles non-smooth penalties. Extensive simulation studies evaluate performance under varying sample sizes, correlation structures, and signal-to-noise ratios, demonstrating that regularized methods (Lasso and Elastic Net) successfully break non-identifiability and achieve low estimation error with controlled false positive rates. An application to NHANES data on asthma and lead exposure illustrates the practical utility: the unregularized estimator yields implausible coefficients, whereas regularized estimators produce stable and interpretable models and automatically select the relevant risk transformation. The Labbe plots derived from regularized estimators indicate a protective effect of reducing lead exposure. The proposed framework bridges identifiability theory with penalized estimation and opens the door to high-dimensional and longitudinal extensions.