🤖 AI Summary
This study addresses the challenges of estimating the number of latent factors and learning the mapping structure in nonlinear latent factor models. Methodologically, it constructs a graph representation based on pairwise dependencies among observed variables and proposes a dependency thresholding algorithm to jointly identify the number of latent factors and the nonlinear mapping structure. A structure-constrained neural network is further designed for efficient optimization. The core contribution lies in overcoming traditional linearity assumptions by establishing an identifiability and consistency theoretical framework grounded in general dependence measures. Simulation experiments demonstrate that the proposed algorithm achieves superior accuracy and robustness in high-dimensional settings, effectively recovering the underlying nonlinear functions.
📝 Abstract
Learning the structure of latent factor models involves two central challenges: (1) estimating the number of latent factors and (2) learning the support of the mapping from latent variables to observed variables. This is especially challenging for nonparametric regimes and nonlinear settings. We propose a method for latent structure learning, based on pairwise dependence measures on the observed variables using a graph-theoretic representation. We show that both the number of latent factors and nonlinear mapping structure can be identified from the distribution of observed variables under mild structural assumptions. Unlike prior work restricted to linear correlations, we establish identifiability and consistency for a general class of dependence measures under nonlinear factor models. This motivates a Dependence Thresholding (DT) algorithm, which jointly estimates the number of latent factors and nonlinear mapping structure from observational data alone. We pair this with a neural network architecture constrained by the nonlinear mapping structure, to recover the nonlinear function. Through simulation studies, we show that the DT algorithm is accurate in practice, even when using flexible methods such as neural networks, and exhibits robustness against violations of its assumptions in high-dimensional settings.