The Semantic Architect: How FEAML Bridges Structured Data and LLMs for Multi-Label Tasks

📅 2025-12-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing LLM-driven feature engineering methods are not designed for multi-label learning, thus failing to model label dependencies and lacking task-specificity. To address this, we propose FEAML—a novel framework that pioneers the integration of LLM-based code generation into multi-label settings. FEAML automatically constructs highly discriminative features by jointly leveraging metadata and label co-occurrence matrices. It introduces label-dependency-aware prompt engineering and a Pearson correlation-based redundancy detection mechanism, coupled with closed-loop optimization guided by classification accuracy. This yields an interpretable, low-redundancy, and self-optimizing feature generation paradigm. Extensive experiments on multiple standard multi-label benchmark datasets demonstrate that FEAML significantly outperforms conventional feature engineering approaches, achieving substantial average improvements in classification accuracy—thereby validating its effectiveness and generalizability.

Technology Category

Machine Learning: Feature Construction/ReformulationNatural Language Processing: GenerationSearch and Optimization: Metareasoning and Metaheuristics

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAG
📝 Abstract
Existing feature engineering methods based on large language models (LLMs) have not yet been applied to multi-label learning tasks. They lack the ability to model complex label dependencies and are not specifically adapted to the characteristics of multi-label tasks. To address the above issues, we propose Feature Engineering Automation for Multi-Label Learning (FEAML), an automated feature engineering method for multi-label classification which leverages the code generation capabilities of LLMs. By utilizing metadata and label co-occurrence matrices, LLMs are guided to understand the relationships between data features and task objectives, based on which high-quality features are generated. The newly generated features are evaluated in terms of model accuracy to assess their effectiveness, while Pearson correlation coefficients are used to detect redundancy. FEAML further incorporates the evaluation results as feedback to drive LLMs to continuously optimize code generation in subsequent iterations. By integrating LLMs with a feedback mechanism, FEAML realizes an efficient, interpretable and self-improving feature engineering paradigm. Empirical results on various multi-label datasets demonstrate that our FEAML outperforms other feature engineering methods.
Problem

Research questions and friction points this paper is trying to address.

FEAML automates feature engineering for multi-label classification tasks.
It models label dependencies using metadata and co-occurrence matrices.
The method integrates feedback to optimize LLM-generated features iteratively.
Innovation

Methods, ideas, or system contributions that make the work stand out.

FEAML automates feature engineering for multi-label classification
It uses LLMs with metadata and label co-occurrence guidance
Incorporates feedback for iterative self-improvement and redundancy detection
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
W
Wanfu Gao
College of Computer Science and Technology, Jilin University, China
Z
Zebin He
College of Computer Science and Technology, Jilin University, China
J
Jun Gao
College of Computer Science and Technology, Jilin University, China