Score
Design and pretrain encoder models that convert heterogeneous table rows into reusable row embeddings via self-supervised objectives, without conditioning on any specific downstream target. Build the encoder architectures, pretraining tasks, and data pipelines to leverage unlabeled tabular corpora so the resulting task-agnostic representations can be applied or fine-tuned across multiple tabular prediction or retrieval tasks (including in-context example conditioning).
This survey addresses fundamental challenges in tabular data representation learning—namely, heterogeneous feature coupling, sample sparsity, and weak task generalization—hindering effective modeling of structured tabular information with deep neural networks. To tackle these issues, we propose a three-tiered methodological framework—*specialized*, *transferable*, and *general-purpose*—and introduce the first taxonomy categorizing models along three dimensions: features, samples, and prediction objectives. We formally define and classify transferable models and tabular foundation models, and unify recent advances in multimodal alignment, open-world adaptation, and self-supervised pretraining. Our work yields the first structured, extensible landscape of tabular representation learning, accompanied by an open-source repository (GitHub) containing curated resources, benchmarks, and implementation guidelines. This synthesis provides both theoretical foundations and practical blueprints for algorithm design and industrial deployment. (138 words)
This work investigates the transferability of task-agnostic embeddings for tabular data, systematically evaluating generative foundation models (e.g., TabPFN, TabICL) against traditional feature engineering—specifically TableVectorizer—on downstream tasks including anomaly detection and supervised learning. The study finds that TableVectorizer, leveraging lightweight, hand-crafted feature transformations, produces high-quality universal embeddings that match or exceed the performance of large foundation models across multiple benchmarks, while achieving inference speeds three orders of magnitude faster. These results challenge the prevailing assumption that complex foundation models are inherently superior for generic tabular representation learning. Crucially, this is the first empirical demonstration that simple, interpretable, and computationally efficient feature engineering can be highly competitive in learning task-agnostic tabular embeddings. The work thus establishes a more practical, deployable paradigm for tabular representation learning—grounded in transparency, efficiency, and robust generalization.
To address the challenge of cross-dataset heterogeneity in tabular data that impedes effective pretraining, this paper proposes TabPTM—a novel framework that maps instances into a shared neighborhood embedding space and constructs meta-representations based on neighbor distances and labels, thereby unifying diverse tabular tasks into homogeneous local prediction problems. TabPTM introduces the first neighborhood-relation-based meta-representation paradigm, enabling zero-shot, cross-dataset transfer without fine-tuning. Technically, it integrates neighborhood embedding, unsupervised distance modeling, and lightweight graph-structured encoding. Extensive experiments across 101 real-world tabular datasets demonstrate that TabPTM significantly outperforms existing pretraining methods on both classification and regression tasks. Notably, it achieves the first truly fine-tuning-free generalization for tabular models, establishing a new benchmark for universal, plug-and-play tabular representation learning.
Existing table embedding methods lack a unified benchmark, making effective cross-task and cross-domain comparisons challenging. To address this gap, this work proposes TEmBed—the first comprehensive evaluation framework for table embeddings that systematically covers four representation granularities: cell, row, column, and full-table levels. Under a consistent experimental setup, the study conducts a systematic assessment of multiple representative models across diverse tasks and domains. The empirical results reveal significant performance variations among existing approaches at different granularities and tasks, offering practical guidance for model selection in real-world applications and advancing the development of general-purpose table representation learning.
Deep learning modeling for tabular data remains challenging due to irregular schema structures and heterogeneous, sparse information distributions. Method: This paper systematically surveys 127 top-tier conference and journal papers published since 2020, proposing a unified analytical framework for tabular representation learning grounded in three pillars: training data, neural architecture, and learning objective. Contribution/Results: It introduces the first holistic “triadic” analysis paradigm, emphasizing cross-task generalizability and robustness. The framework comprehensively covers key advances—including data augmentation, specialized architectures (e.g., FT-Transformer), self-supervised pretraining, contrastive learning, and multi-task optimization. It identifies critical research gaps, distills fundamental challenges and evolutionary trends, and establishes standardized evaluation dimensions. Collectively, this work provides both theoretical foundations and practical guidelines for developing general-purpose, robust, and interpretable deep learning methods for tabular data.
This work challenges the prevailing assumption that Table Foundation Models (TFMs) require large-scale synthetic or real-world pretraining data to achieve generalization. Method: We propose a lightweight self-supervised pretraining framework that learns structured semantics from a single real-world table, combined with in-context learning for zero-shot cross-domain transfer—without external corpora or additional annotations. Contribution/Results: We demonstrate that the quality and diversity of task construction—not data scale—are the primary determinants of TFM performance. Evaluated across heterogeneous downstream benchmarks spanning finance, healthcare, and e-commerce, our approach significantly outperforms existing few-shot baselines. These results validate the effectiveness and scalability of the “single-table pretraining + in-context learning” paradigm, establishing a novel, resource-efficient framework for tabular modeling in low-data regimes.
This work addresses the lack of systematic evaluation of table-level embeddings in multi-task settings, which hinders a comprehensive assessment of their downstream effectiveness. To overcome the limitations of single-task benchmarks focused solely on retrieval, we introduce TEmBed-T—a multidimensional benchmark encompassing tasks such as retrieval, classification, and data lake discovery. TEmBed-T integrates diverse table embedding models and establishes a unified evaluation protocol with consistent metrics across tasks. Empirical results demonstrate that no single embedding model consistently outperforms others across all tasks, revealing that the quality of table-level embeddings must be evaluated holistically through multiple dimensions rather than relying exclusively on retrieval performance.
Existing tabular in-context learning methods often couple feature representations with specific prediction targets, limiting their generalization across diverse tasks. This work proposes a task-agnostic encoder-decoder architecture that employs a single shared encoder to learn universal row embeddings from unlabeled real-world tables, paired with multiple task-specific decoders to support six downstream tasks: classification, regression, anomaly detection, clustering, entity matching, and entity classification in relational databases. To the best of our knowledge, this is the first approach to enable effective transfer of a unified representation across multiple tabular in-context learning tasks. The method achieves state-of-the-art performance on most tasks and substantially enhances model generalization.
研究探讨了表格基础模型通过单一真实表格自监督预训练实现强迁移的问题,提出任务中心和基于检索的视角来解释其泛化性能。
Existing table encoders are difficult to compare fairly in terms of representational capacity due to their evaluation within task-specific end-to-end pipelines. This work proposes TRL-Bench, a multi-granularity benchmark for table representation learning, which standardizes evaluation across row-, column-, and table-level granularities by providing a unified interface to export embeddings, lightweight shared probing heads, and three evaluation suites—TRL-CTBench, TRL-RBench, and TRL-DLTE—covering 16 tasks across 20 models. For the first time, this framework enables comparable assessment of encoder representations across diverse training paradigms, revealing pronounced task specificity: task-aligned specialized encoders consistently outperform general-purpose models, and optimal end-to-end performance hinges on non-additive, synergistic compatibility among pipeline components rather than the ranking of any single stage.
This work addresses the challenge that existing tabular foundation models, such as TabPFN, struggle to natively handle high-cardinality textual features, often resorting to PCA-based compression of text embeddings—a process prone to significant information loss. To overcome this limitation, the authors propose a lightweight text adapter that maps frozen sentence encoder outputs into short token sequences residing in TabPFN’s embedding space. This approach enables efficient fusion of textual and tabular data without requiring end-to-end retraining. Inspired by cross-modal projection techniques, the method preserves TabPFN’s strong numerical modeling capabilities while circumventing the PCA bottleneck, leading to substantial performance gains on tabular tasks involving textual features.