🤖 AI Summary
This study addresses the failure of objective function minimization along signal-free directions in projection pursuit, caused by Gaussian complements, by proposing to restrict the search to the column space of a known operator. Through analyzing the effects of sample splitting and coordinate augmentation, we reveal the gain discrepancies between covariance spike and fourth-moment searches, establishing empirical scaling laws under fixed search dimensions. Numerical simulations based on principal component estimation, projection pursuit algorithms, and controlled two-component models demonstrate that the measured threshold ratio exponentially approaches the theoretically predicted value of 1/8; however, this scaling effect is substantially attenuated under downstream error criteria.
📝 Abstract
Projection pursuit searches for a direction along which the data look least Gaussian. When the observation space contains a large Gaussian complement, the empirical objective can be minimized by a direction that carries no signal, with empirical kurtosis as low as at the truth. Sample splitting exposes rather than repairs this failure. Appending coordinates independent of the latent regime degrades the search while leaving Bayes recoverability unchanged. Restricting the search to the column space of a known forward operator removes the failure exactly on the negative-kurtosis branch. Estimating a principal subspace from the data is the alternative. In a controlled two-component model, the leading sufficient scalings differ in the gain with which the operator transmits the discriminant: $ς^{-4}$ for covariance-spike estimation and $ς^{-8}$ for fourth-moment search. At fixed search dimension, the measured threshold ratio collapses onto $n/p^2$ with exponent $0.156$, close to the predicted $1/8$. This is an empirically supported scaling motivated by sufficient bounds, not a proved asymptotically tight law. When the search dimension is varied, the measured exponent is $0.325$, substantially larger than $1/8$, and the tested range does not identify its functional form. The crossing location also depends on calibration and model configuration. Under a downstream excess-error criterion, the scaling largely disappears.