LLMs as Feature Engineers for Text-and-Tabular Prediction

📅 2026-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种迭代框架,利用生成器和提取器LLM从非结构化文本中自动提取可解释的分类特征,并通过下游模型评估其预测性能,以加速表格预测模型的特征发现。
📝 Abstract
We introduce an iterative framework that automates the extraction of interpretable, schema-bound categorical features from unstructured text for tabular prediction models. To navigate the feature space, a generator LLM proposes semantic definitions, a separate extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance. We optimize this search by translating explicit model errors, such as AUC ranking inversions, into natural-language feedback, steering the LLM to resolve specific predictive failures. Evaluated across three public datasets, this error-driven loop accelerates feature discovery by up to $3\times$ compared to unguided search. Empirically, the generated features demonstrate strong multi-view complementarity, strictly outperforming any subset when combined with TF-IDF and dense embeddings. Finally, the framework guarantees instance-level interpretability: the discovered features dominate SHAP importance rankings and provide a fully transparent, semantic audit trail for every prediction.
Problem

Research questions and friction points this paper is trying to address.

LLM
feature extraction
tabular prediction
unstructured text
predictive performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

iterative framework
large language models
feature extraction
error-driven optimization
multi-view complementarity