Rethinking Pre-Training in Tabular Data: A Neighborhood Embedding Perspective

πŸ“… 2023-10-31
πŸ“ˆ Citations: 6
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
To address the challenge of cross-dataset heterogeneity in tabular data that impedes effective pretraining, this paper proposes TabPTMβ€”a novel framework that maps instances into a shared neighborhood embedding space and constructs meta-representations based on neighbor distances and labels, thereby unifying diverse tabular tasks into homogeneous local prediction problems. TabPTM introduces the first neighborhood-relation-based meta-representation paradigm, enabling zero-shot, cross-dataset transfer without fine-tuning. Technically, it integrates neighborhood embedding, unsupervised distance modeling, and lightweight graph-structured encoding. Extensive experiments across 101 real-world tabular datasets demonstrate that TabPTM significantly outperforms existing pretraining methods on both classification and regression tasks. Notably, it achieves the first truly fine-tuning-free generalization for tabular models, establishing a new benchmark for universal, plug-and-play tabular representation learning.
πŸ“ Abstract
Pre-training is prevalent in deep learning for vision and text data, leveraging knowledge from other datasets to enhance downstream tasks. However, for tabular data, the inherent heterogeneity in attribute and label spaces across datasets complicates the learning of shareable knowledge. We propose Tabular data Pre-Training via Meta-representation (TabPTM), aiming to pre-train a general tabular model over diverse datasets. The core idea is to embed data instances into a shared feature space, where each instance is represented by its distance to a fixed number of nearest neighbors and their labels. This ''meta-representation'' transforms heterogeneous tasks into homogeneous local prediction problems, enabling the model to infer labels (or scores for each label) based on neighborhood information. As a result, the pre-trained TabPTM can be applied directly to new datasets, regardless of their diverse attributes and labels, without further fine-tuning. Extensive experiments on 101 datasets confirm TabPTM's effectiveness in both classification and regression tasks, with and without fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Pre-training for heterogeneous tabular data
Shared feature space via meta-representation
General tabular model without fine-tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Meta-representation for tabular data
Neighborhood embedding feature space
General model without fine-tuning
Nanjing University
Han-Jia Ye
Han-Jia Ye
Nanjing University
Machine LearningData MiningMetric LearningMeta-Learning
Q
Qi-Le Zhou
National Key Laboratory for Novel Software Technology, Nanjing University
De-Chuan Zhan
De-Chuan Zhan
Nanjing University, China
Machine LearningData Mining
H
Huai-Hong Yin
W
Wei-Lun Chao