๐ค AI Summary
In federated learning (FL), data silos impede the application of automated feature engineering (AutoFE), leading to suboptimal model performance due to the absence of coordinated, high-quality feature construction across clients.
Method: This paper proposes the first privacy-preserving AutoFE framework compatible with horizontal, vertical, and hybrid FL settings. Without requiring raw data to leave local domains, it jointly constructs valuable features via federated optimization, distributed feature generation, differential privacy- or secure aggregationโdriven feature fusion, and cross-client feature-space alignment.
Contribution/Results: Experiments on multiple FL benchmark tasks demonstrate that the method improves downstream model AUC by 3.2โ5.7% over baseline approaches, closely approaching centralized AutoFE performance and significantly outperforming isolated client-side feature engineering. It establishes the first systematic study of AutoFE in federated environments, bridging a critical gap in privacy-aware machine learning.
๐ Abstract
Automated feature engineering (AutoFE) is used to automatically create new features from original features to improve predictive performance without needing significant human intervention and domain expertise. Many algorithms exist for AutoFE, but very few approaches exist for the federated learning (FL) setting where data is gathered across many clients and is not shared between clients or a central server. We introduce AutoFE algorithms for the horizontal, vertical, and hybrid FL settings, which differ in how the data is gathered across clients. To the best of our knowledge, we are the first to develop AutoFE algorithms for the horizontal and hybrid FL cases, and we show that the downstream test scores of our federated AutoFE algorithms is close in performance to the case where data is held centrally and AutoFE is performed centrally.