latent factor ensembling

Design and build ensembles of latent-factor models (for example alternating-least-squares matrix factorizations) that learn compact, largely time-invariant latent features by minimizing reconstruction error via alternating updates. Combine diverse factorization models to fuse complementary latent representations, produce more comprehensive and robust embeddings, and scale the factorization to large or sparse datasets.

latentfactorensembling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.8
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing latent factor models in handling high-dimensional incomplete (HDI) data, which often yield biased and insufficient representations due to their sole reliance on gradient descent optimization. To overcome this issue, the authors propose a novel heterogeneous ensemble approach that uniquely integrates differential evolution and gradient descent to construct two complementary latent factor models. An adaptive weighting mechanism is further introduced to dynamically fuse the strengths of both models during training. This strategy effectively mitigates representation bias and consistently outperforms several state-of-the-art latent factor models across three HDI datasets, demonstrating its superior effectiveness and robustness.

High-dimensional and Incomplete DataLatent Factor ModelOptimization Limitation

Transfer learning for high-dimensional Factor-augmented sparse model

Nov 15, 2025
BF
Bo Fu
🏛️ Xi'an Jiaotong University

In economics and finance, sparse linear models often exhibit high dimensionality, strong correlations, and latent factor structures; yet target samples are scarce and highly susceptible to model misspecification, rendering conventional estimators unreliable. Method: We propose a transfer learning framework leveraging heterogeneous auxiliary data from multiple sources. First, we design a data-driven source detection algorithm to automatically identify informative auxiliary datasets and mitigate negative transfer. Second, we develop a hypothesis testing framework for factor model applicability and enable joint confidence interval inference for regression coefficients. Third, via non-asymptotic analysis, we derive ℓ₁/ℓ₂ estimation error bounds to ensure robustness. Contribution/Results: Theoretically, our estimator achieves the optimal convergence rate. Simulation studies and empirical applications demonstrate substantial improvements in estimation accuracy and statistical reliability under cross-dataset heterogeneity.

Addresses high-dimensional factor-augmented sparse model estimation challengesDevelops transfer learning procedures leveraging heterogeneous auxiliary datasetsMitigates predictor correlation and latent factor misspecification in linear models

This work addresses the challenge of effectively integrating user–item and item–item collaborative filtering to enhance Top-N recommendation performance while maintaining computational efficiency. The authors propose a weighted similarity ensemble method based on shared embeddings, which, for the first time, unifies both recommendation pathways within a single framework. By sharing user and item embeddings across strategies, the approach simplifies model architecture and eliminates the need for separate hyperparameter tuning for each pathway, thereby substantially reducing deployment complexity. Experimental results demonstrate that the proposed method achieves competitive recommendation accuracy across multiple datasets and exhibits robust performance in scenarios favoring different collaborative filtering paradigms.

Collaborative FilteringItem-Item SimilarityRecommender Systems

Optimal discriminant analysis in high-dimensional latent factor models

Oct 23, 2022
XB
Xin Bing
🏛️ University of Toronto | Cornell University

This paper addresses classification under high-dimensional sparse settings. We propose a two-step discriminant method based on principal component analysis (PCA), grounded in an implicit low-rank factor model and featuring adaptive selection of the number of principal components. We establish, for the first time, a general risk analysis framework for high-dimensional two-step classifiers and rigorously derive the minimax-optimal convergence rate (up to logarithmic factors) for the PCA-based classifier—even when dimensionality far exceeds sample size. Theoretically, the excess risk achieves the optimal rate; simulations demonstrate robustness under model misspecification; and empirical evaluation on three real-world high-dimensional datasets shows significant improvement over state-of-the-art discriminant methods. Key contributions include: (i) a unified theoretical analysis paradigm for two-step classification, (ii) minimax-optimal rate guarantees, (iii) a data-driven, theoretically justified dimension-selection mechanism, and (iv) consistent empirical superiority across diverse high-dimensional benchmarks.

Analyze convergence rates of excess risk in classificationDevelop efficient classifier for high-dimensional latent factor modelsSelect optimal principal components in data-driven projection

Linear combinations of latents in generative models: subspaces and beyond

Aug 16, 2024
EB
Erik Bodin
🏛️ University of Cambridge | Lancaster University

To address the lack of universality and interpretability in latent variable manipulation for generative models, this paper proposes Linear Latent Composition (LOL). LOL establishes the first modality-agnostic and architecture-agnostic framework enabling arbitrary linear operations—including interpolation, subspace construction, and low-dimensional representation extraction—without additional training or fine-tuning. Its core innovation lies in applying geometrically consistent linear transformations within the latent spaces of mainstream generative paradigms, including diffusion models, flow matching, and continuous normalizing flows. Unlike existing approaches constrained to specific architectures or data modalities, LOL significantly enhances the flexibility and reproducibility of controllable generation. Empirical evaluations demonstrate its strong generalization and practical utility across diverse applications: synthetic data generation, data augmentation, and multimodal experimental design.

Control over generative model latent variablesGeneral-purpose method for linear combinationsSimplifies creation of low-dimensional representations

Latest Papers

What's happening recently
View more

Identifying latent factors that are stable and transferable across distributionally heterogeneous environments is a key challenge for robust cross-environment prediction. This work proposes ATLAS, the first method to strictly disentangle invariant from environment-specific factors via structural conditioning in heterogeneous settings with partially available auxiliary labels, while unifying treatment of both labeled and unlabeled scenarios for transfer prediction. ATLAS integrates invariance-guided latent alignment, exploitation of auxiliary supervision signals, and low-dimensional latent factor regression, and comes with non-asymptotic theoretical guarantees on prediction error. Experiments demonstrate that ATLAS nearly perfectly recovers the true invariant factors and achieves oracle-level regression performance by enabling full latent signal transfer to novel environments.

heterogeneous environmentsinvariant factorslatent factor model

This study addresses the challenging problem of recovering a low-rank latent factor structure linked through an unknown monotonic nonlinear function from incomplete and noisy observations, which is hindered by severe non-convexity and identifiability ambiguities. The authors generalize linear factor models to a nonparametric nonlinear setting by assuming the link function resides in a reproducing kernel Hilbert space (RKHS) and introduce explicit regularization to resolve scale and rotational indeterminacies. They propose a projected block coordinate descent algorithm that jointly estimates the latent factors, loading matrix, and link function, providing theoretical convergence guarantees in both noiseless and noisy regimes. Moreover, the update of the link function enjoys a sublinear regret bound. Synthetic experiments demonstrate the method’s effectiveness and robustness.

identifiabilityincomplete datanoisy data

This study addresses the inherent non-identifiability of latent variables in factor models—manifested as non-uniqueness and distributional shifts—by systematically elucidating their nature in linear factor models and their implications for representation learning, drawing on an interdisciplinary perspective spanning psychometrics, statistics, and artificial intelligence. It establishes a theoretical connection between this identifiability issue and posterior collapse in variational autoencoders. By integrating factor analysis, linear autoencoders, and variational inference within a high-dimensional asymptotic framework, the work proves that latent factors become fully identifiable as the observation dimension tends to infinity. Building on this result, the authors propose a nearly distribution-free estimation method for high-dimensional settings, effectively bridging the theoretical gap between classical factor analysis and modern deep generative models, particularly well-suited for representation learning with ultra-high-dimensional data.

data representationfactor analysisgenerative models

This study addresses the joint tri-factorization of symmetric multi-type nonnegative matrices, aiming to learn shared orthogonal nonnegative factors that facilitate interpretable clustering and network analysis. It introduces, for the first time, a shared orthogonal nonnegative factor structure that enhances interpretability while preserving hard assignment properties. To optimize this formulation, two efficient algorithms are proposed: one based on a penalty-function fixed-point method derived from KKT conditions, and another employing a three-stage strategy that integrates nonnegative optimization, orthogonalization, and constrained ADAM fine-tuning. Experimental results demonstrate that the proposed methods reliably recover near-optimal factorizations on noisy synthetic data and yield embeddings that match or outperform established baselines—including SVD and node2vec—in citation network benchmarks across link prediction, node clustering, and classification tasks.

clusteringnetwork analysisnon-negativity

This work addresses a critical limitation in existing factorized generative models, which only match the marginal distribution of style latent variables without enforcing independence from class information, leading to conditional style leakage. The study demonstrates for the first time that marginal distribution matching alone is insufficient for effective disentanglement and establishes that four theoretical conditions must be jointly satisfied. To systematically quantify style-class leakage, the authors introduce a comprehensive auditing framework combining maximum mean discrepancy (MMD), linear probing, clustering evaluation, and multidimensional perturbation experiments. Empirical results reveal that multiple baseline models—despite achieving near-zero marginal MMD—still enable label recovery with 74%–100% accuracy. The proposed post-processing method substantially improves generation quality, attaining external evaluation scores of 0.97 on MNIST and 0.88 on CIFAR-10.

class-conditional generationconditional style leakagefactorized generative models

Hot Scholars

KK

Klemen Kotar

PhD Candidate, Stanford University
Artificial Intelligence
YM

Yong Ma

Wuhan University
Infrared image processingremote sensing
ML

Miguel Lastra

Profesor Contratado Doctor de la Universidad de Granada (Associate Professor)
GPU programmingMachine LearningTime SeriesLarge Scale Biometric Systems
YP

Yan Pang

University of Colorado
Computer VisionMedical Image AnalysisGraph Neural Network
BX

Bin Xiao

Meta GenAI
Computer VisionVision and LanguageMachine LearningHuman Pose Estimation