learn tabular embeddings

Designs and trains models that produce dense vector embeddings for tabular data columns and records, including contextual categorical embeddings and TabTransformer-style architectures. These embeddings capture cross-feature semantics and interactions (for example via self-attention) and fuse continuous and categorical inputs to yield compact latent representations for downstream prediction, ranking, or retrieval tasks.

learntabularembeddings

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$208K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting

Feb 20, 2025
GS
Gyanendra Shrestha
🏛️ Florida State University | University of South Florida

This work addresses the challenge of modeling complex two-dimensional contextual relationships in 2D tables featuring bidirectional hierarchical metadata and nested structures. We propose the first embedding representation method specifically designed for such tables. Our approach introduces three key innovations: (1) bidimensional coordinate encoding to explicitly capture row- and column-wise hierarchical dependencies; (2) a hierarchical visibility matrix that decouples metadata context from data cell context; and (3) a structure-aware embedding learning framework with end-to-end nested structure modeling. By overcoming the limitations of conventional flattened table representations, our method achieves significant improvements over state-of-the-art baselines across five large-scale structured datasets and three downstream tasks—including table retrieval, question answering, and schema matching—with up to +0.28 mean average precision (MAP). Notably, it outperforms GPT-4 augmented with retrieval-augmented generation (RAG) by +0.42 MAP.

Encode horizontal and vertical hierarchical metadata.Improve performance on structured dataset tasks.Optimize embeddings for 2-D table structures.

Foundation models have established unified representations for natural language processing, yet this paradigm remains largely unexplored for tabular data. Existing methods face fundamental limitations: LLM-based approaches lack retrieval-compatible vector outputs, whereas text embedding models often fail to capture tabular structure and numerical semantics. To bridge this gap, we first introduce the Tabular Embedding Benchmark (TabBench), a comprehensive suite designed to evaluate the tabular understanding capability of embedding models. We then propose TabEmbed, the first generalist embedding model that unifies tabular classification and retrieval within a shared embedding space. By reformulating diverse tabular tasks as semantic matching problems, TabEmbed leverages large-scale contrastive learning with positive-aware hard negative mining to discern fine-grained structural and numerical nuances. Experimental results on TabBench demonstrate that TabEmbed significantly outperforms state-of-the-art text embedding models, establishing a new baseline for universal tabular representation learning. Code and datasets are publicly available at https://github.com/qiangminjie27/TabEmbed and https://huggingface.co/datasets/qiangminjie27/TabBench.

embedding modelsnumerical semanticsretrieval

Existing table embedding methods lack a unified benchmark, making effective cross-task and cross-domain comparisons challenging. To address this gap, this work proposes TEmBed—the first comprehensive evaluation framework for table embeddings that systematically covers four representation granularities: cell, row, column, and full-table levels. Under a consistent experimental setup, the study conducts a systematic assessment of multiple representative models across diverse tasks and domains. The empirical results reveal significant performance variations among existing approaches at different granularities and tasks, offering practical guidance for model selection in real-world applications and advancing the development of general-purpose table representation learning.

benchmarkfoundation modelsrepresentation learning

To address the challenge of cross-dataset heterogeneity in tabular data that impedes effective pretraining, this paper proposes TabPTM—a novel framework that maps instances into a shared neighborhood embedding space and constructs meta-representations based on neighbor distances and labels, thereby unifying diverse tabular tasks into homogeneous local prediction problems. TabPTM introduces the first neighborhood-relation-based meta-representation paradigm, enabling zero-shot, cross-dataset transfer without fine-tuning. Technically, it integrates neighborhood embedding, unsupervised distance modeling, and lightweight graph-structured encoding. Extensive experiments across 101 real-world tabular datasets demonstrate that TabPTM significantly outperforms existing pretraining methods on both classification and regression tasks. Notably, it achieves the first truly fine-tuning-free generalization for tabular models, establishing a new benchmark for universal, plug-and-play tabular representation learning.

General tabular model without fine-tuningPre-training for heterogeneous tabular dataShared feature space via meta-representation

The evolution of embedding techniques from word vectors to multimodal representations remains fragmented, lacking a unified framework that integrates advances across linguistic, cross-lingual, personalized, and multimodal domains—particularly for embodied multimodal learning in large language models. Method: We systematically survey static and contextual language representations, cross-lingual and personalized modeling, sentence/document embeddings, and multimodal fusion in vision, robotics, and cognitive science. We synthesize recent progress in interpretability, model compression, numerical encoding, and bias mitigation, and propose a novel paradigm emphasizing strong alignment across non-textual modalities and scalable training. Contributions: We construct a comprehensive knowledge graph of end-to-end embedding technologies—from Word2Vec and BERT to GPT, generative topic models, and multimodal alignment/distillation methods—identifying key technical bottlenecks and ethical challenges. This work delivers the first systematic roadmap for multimodal, embodied learning in foundation models.

Addressing compression, interpretability and bias challengesEvolving from sparse to dense word embeddingsExtending embeddings to multimodal domains

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic evaluation of table-level embeddings in multi-task settings, which hinders a comprehensive assessment of their downstream effectiveness. To overcome the limitations of single-task benchmarks focused solely on retrieval, we introduce TEmBed-T—a multidimensional benchmark encompassing tasks such as retrieval, classification, and data lake discovery. TEmBed-T integrates diverse table embedding models and establishes a unified evaluation protocol with consistent metrics across tasks. Empirical results demonstrate that no single embedding model consistently outperforms others across all tasks, revealing that the quality of table-level embeddings must be evaluated holistically through multiple dimensions rather than relying exclusively on retrieval performance.

benchmarkdownstream taskssystematic evaluation

Existing table annotation methods linearize two-dimensional tables into one-dimensional sequences, which compromises semantic expressiveness, discards structural information, and limits model generalization. This work proposes TabEmb, the first approach to decouple semantic encoding from structural modeling: it leverages large language models to generate column-wise semantic embeddings and explicitly captures inter-column relationships using a graph neural network, thereby enabling joint representation learning of semantics and structure. Evaluated across multiple table annotation tasks, TabEmb substantially outperforms strong existing baselines, significantly improving both annotation quality and model generalization.

2D-to-1D flatteningcontext-length constraintssemantic embedding

Existing tabular in-context learning methods often couple feature representations with specific prediction targets, limiting their generalization across diverse tasks. This work proposes a task-agnostic encoder-decoder architecture that employs a single shared encoder to learn universal row embeddings from unlabeled real-world tables, paired with multiple task-specific decoders to support six downstream tasks: classification, regression, anomaly detection, clustering, entity matching, and entity classification in relational databases. To the best of our knowledge, this is the first approach to enable effective transfer of a unified representation across multiple tabular in-context learning tasks. The method achieves state-of-the-art performance on most tasks and substantially enhances model generalization.

downstream tasksencoder-decoder architecturein-context learning

This work addresses the limitations of existing tabular learning methods, which treat feature names and cell values as discrete symbols and thus overlook their semantic content, leading to degraded performance under data sparsity. To overcome this, the authors propose CASE, a framework that integrates the semantic comprehension of large language models with the statistical modeling capacity of tabular learners through a context-aware mechanism for row embedding generation. The core innovation lies in a customized Gemma-3-based tabular language model augmented with a pre-filled KV cache strategy, which leverages representative samples as semantic anchors to produce dynamically contextualized embeddings, effectively mitigating semantic ambiguity. Experimental results demonstrate that CASE achieves substantial performance gains across multiple benchmarks—including CARTE, TextTab, and TabArena—with particularly pronounced advantages in low-resource settings.

context-awarenessfeature semanticsLLM integration

This work addresses the challenge that existing tabular foundation models, such as TabPFN, struggle to natively handle high-cardinality textual features, often resorting to PCA-based compression of text embeddings—a process prone to significant information loss. To overcome this limitation, the authors propose a lightweight text adapter that maps frozen sentence encoder outputs into short token sequences residing in TabPFN’s embedding space. This approach enables efficient fusion of textual and tabular data without requiring end-to-end retraining. Inspired by cross-modal projection techniques, the method preserves TabPFN’s strong numerical modeling capabilities while circumventing the PCA bottleneck, leading to substantial performance gains on tabular tasks involving textual features.

high-cardinality text featuresinformation bottleneckPCA compression

Hot Scholars

SD

Subasish Das

Assistant Professor, Texas State University and GBD Senior Collaborator
Civil EngineeringSafetyTransportation EngineeringAI
JS

Jieming Shi

The Hong Kong Polytechnic University
Data ManagementData miningBig data analytics
SS

Shriyank Somvanshi

Doctoral Student, Texas State University
Deep LearningTabular DataScientific Machine LearningTransportation Safety
LP

Li Pan

Shanghai Jiao Tong University