feature engineering

Designs and builds transformations and selection pipelines that convert raw inputs—such as cepstral/audio, textual, log, sensor/time-series, sequence, node/relational signals—into model-ready features. Work covers hand-crafted and automated approaches (including program-synthesized or evolutionary/LLM-guided python feature programs), normalization and combination of signals, redundancy removal and selection, and empirical validation/optimization of features for downstream model performance.

featureengineering

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-1.25
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$196K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the automatic discovery of efficient and interpretable preprocessing and feature engineering methods for structured data, such as time series and tabular datasets. It proposes an Evolutionary Feature Engineering (EFE) framework that, for the first time, employs large language models as evolutionary search operators to automatically generate Python programs conforming to the standard fit/transform interface. The framework iteratively optimizes these programs by leveraging data context, statistical summaries, and validation performance, enabling end-to-end compatibility with existing machine learning pipelines. On time series tasks, the approach reduces MASE, WQL, and MAE errors by over 3% on average (up to 19%). For tabular data, EFE-Tab produces compact feature sets that achieve competitive accuracy with decision tree models while preserving strong interpretability.

Evolutionary OptimizationFeature EngineeringInterpretability

Dynamic and Adaptive Feature Generation with LLM

Jun 04, 2024
XZ
XinHao Zhang
🏛️ Portland State University | Chinese Academy of Sciences | University of Chinese Academy of Sciences

Existing feature engineering approaches suffer from three fundamental limitations: poor interpretability, weak generalizability, and inflexible strategies—hindering practical deployment across diverse scenarios. To address these challenges, this paper proposes the first large language model (LLM)-driven dynamic adaptive feature generation paradigm. Our method integrates task-aware prompting with semantic modeling of the feature space, enabling real-time, interpretable, and controllable feature generation tailored to both data characteristics and task requirements. It ensures cross-modal and cross-task generality while maintaining full transparency in the feature generation process. Extensive experiments on multiple structured and unstructured data tasks demonstrate that features generated by our approach improve feature quality by 23.6% and boost downstream model performance by an average of 11.4%, significantly outperforming conventional automated feature engineering methods.

Enhances applicability across diverse data typesImproves explainability of feature generation processIncreases strategic flexibility in feature engineering

Machine learning model selection lacks formalized methodologies, making it difficult to systematically characterize contextual factors—such as data characteristics and prediction tasks—and their interactions, resulting in opaque, non-adaptive decisions. This paper introduces, for the first time, software product line (SPL) principles into ML model selection, proposing a variability-aware algorithm selection framework. It constructs a configurable feature model that explicitly captures commonalities and variabilities among contextual factors—including dataset size, feature dimensionality, and task type—as well as their logical dependencies. By integrating scikit-learn’s heuristic rules with an instantiation framework, the approach enables interpretable, adaptive, and transparent model recommendations. An empirical case study demonstrates that the method significantly outperforms existing strategies in accuracy, interpretability, and contextual adaptability.

Machine LearningModel SelectionRule Formalization

This work addresses the cumbersome and error-prone nature of data preparation and feature engineering in machine learning pipelines, as well as the limited controllability of current large language model (LLM)-assisted programming approaches, which hinders their deployment in production systems. The authors propose SemPipes, a declarative, LLM-driven semantic operator programming model that seamlessly integrates natural language instructions with Python code and automatically synthesizes dataset- and context-aware implementations during training. By innovatively combining declarative semantic operators with evolutionary search optimization, SemPipes enables controlled and optimizable integration of LLMs into ML pipelines. The accompanying SemPiper interactive interface allows users to edit pipelines, inspect generated code, and observe operator synthesis and optimization in real time, demonstrating the approach’s flexibility and practical utility.

code synthesislarge language modelsmachine learning pipelines

ELATE: Evolutionary Language model for Automated Time-series Engineering

Aug 20, 2025
AM
Andrew Murray
🏛️ JP Morgan AI Research

In time-series forecasting, manual feature engineering is labor-intensive and inefficient, while existing automated approaches often rely on computationally expensive exhaustive search and lack domain awareness. To address these limitations, we propose the first paradigm integrating large language models (LLMs) into an evolutionary feature generation framework. Specifically, the LLM generates semantically meaningful and context-aware candidate transformations grounded in time-series statistics and feature importance; an evolutionary algorithm then efficiently searches and prunes this space, substantially reducing redundant computation. Our method eliminates reliance on hand-crafted rules or brute-force enumeration. Evaluated across diverse benchmarks, it achieves an average 8.4% improvement in forecasting accuracy. This demonstrates that language-guided intelligent feature engineering delivers synergistic gains in interpretability, computational efficiency, and predictive performance.

Automating time-series feature engineering to avoid manual effortIncorporating domain-specific insights into feature transformation processesReducing computational costs of exhaustive enumeration methods

Latest Papers

What's happening recently
View more

This work proposes an end-to-end, feature-free audio classification approach based on a parallel deep reservoir computing architecture that operates directly on raw audio waveforms, eliminating the need for explicit feature extraction such as MFCCs. Traditional methods relying on handcrafted features often incur high computational overhead and complex preprocessing pipelines. To evaluate the efficacy of the proposed design, the authors conduct comparative experiments using shallow, serial, and parallel deep reservoir models. Results demonstrate that the parallel architecture achieves significantly superior performance over baseline methods while maintaining low model complexity. The approach enables efficient temporal modeling and hierarchical representation learning, highlighting its scalability and practical potential for audio processing tasks.

acoustic signal preprocessingend-to-end classificationfeature-free

This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.

code synthesisfluid systemslarge language models

Machine Learning Pipeline for Software Engineering: A Systematic Literature Review

Jul 31, 2025
SK
Samah Kansab
🏛️ École de technologie supérieure (ÉTS)

Addressing the longstanding challenge in software engineering (SE) of balancing quality assurance with development efficiency, this study proposes a systematically derived and optimized machine learning (ML) pipeline framework tailored for SE tasks. Methodologically, the pipeline integrates automated data acquisition, SMOTE-based class balancing, SZZ-inspired feature selection, ensemble models (Random Forest and Gradient Boosting), and a novel evaluation metric—Balanced Accuracy Metric (BAM)—alongside standard metrics (AUC, F1-score, precision) and bootstrap resampling for robust validation. Key contributions include: (1) the first structured, end-to-end optimization framework for SE-specific ML pipelines; (2) empirical evidence demonstrating superior defect prediction performance of ensemble methods over individual classifiers; and (3) identification of two critical research gaps—insufficient data standardization and lack of model interpretability—thereby establishing foundational theoretical insights and actionable directions for intelligent SE research.

Addressing scalability issues in traditional SE approaches with MLEnsuring quality and efficiency in complex software engineering lifecycleOptimizing ML pipelines for robust defect prediction and automation

MLPrE -- A tool for preprocessing and exploratory data analysis prior to machine learning model construction

Oct 29, 2025
DS
David S Maxwell
🏛️ The University of Texas MD Anderson Cancer Center

To address poor scalability, integration complexity, and inflexible configuration in preprocessing multi-source heterogeneous data for machine learning modeling and graph database construction, this paper proposes a lightweight, modular, JSON-driven automated data preprocessing framework. Built upon Spark DataFrames for efficient distributed processing, the framework defines 69 composable and parallelizable processing stages spanning input parsing, filtering, statistical analysis, feature engineering, and graph-structure transformation. Its declarative JSON-based configuration enables dynamic adaptation to varying data types and scales, significantly enhancing interoperability with workflow orchestration systems such as Apache Airflow. Experimental evaluation across six heterogeneous datasets demonstrates the framework’s generality and scalability: it successfully supports wine quality clustering analysis and end-to-end conversion of phosphosylation site–kinase interaction data into a graph database.

Addressing scalability limitations in existing data processing workflowsPreprocessing diverse data formats for machine learning modelsProviding exploratory analysis capabilities for large-scale datasets

This study addresses the limitations of existing approaches for automatically extracting machine learning (ML) pipeline structures, which often rely on manual annotations or suffer from insufficient generalization to keep pace with the rapid evolution of the ML ecosystem. The work presents the first systematic evaluation of small language models (SLMs) for reverse-engineering ML pipelines and proposes an SLM-based method for their automatic identification and reconstruction. Through comprehensive comparative experiments across multiple SLMs and rigorous statistical validation using Cochran’s Q, McNemar, and Pearson’s chi-squared tests, the authors demonstrate that the best-performing SLM significantly outperforms current methods and exhibits robustness across diverse classification schemes. This approach uncovers finer-grained patterns in data science practices and overcomes longstanding bottlenecks in scalability and domain adaptability inherent in traditional techniques.

Code UnderstandingData Science PracticesMachine Learning Pipelines

Hot Scholars

AA

Aditya Akella

Professor, Computer Science, UT Austin
Computer NetworksNetworkingComputer SystemsSystems
HZ

Hongsong Zhu

institute of information Engineering, Chinese Academy of Sciences
cybersecurityinternet measurement
ZH

Zheng He

University of British Columbia
deep learningmachine learning
BM

Baichuan Mo

PhD @ MIT, Research Scientist @ TikTok, Lyft
TransportationOptimizationMachine LearningDemand Modeling
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy