Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting

📅 2025-02-20
🏛️ International Conference on Extending Database Technology
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of modeling complex two-dimensional contextual relationships in 2D tables featuring bidirectional hierarchical metadata and nested structures. We propose the first embedding representation method specifically designed for such tables. Our approach introduces three key innovations: (1) bidimensional coordinate encoding to explicitly capture row- and column-wise hierarchical dependencies; (2) a hierarchical visibility matrix that decouples metadata context from data cell context; and (3) a structure-aware embedding learning framework with end-to-end nested structure modeling. By overcoming the limitations of conventional flattened table representations, our method achieves significant improvements over state-of-the-art baselines across five large-scale structured datasets and three downstream tasks—including table retrieval, question answering, and schema matching—with up to +0.28 mean average precision (MAP). Notably, it outperforms GPT-4 augmented with retrieval-augmented generation (RAG) by +0.42 MAP.

Technology Category

Machine Learning: Structured LearningComputer Vision: Representation Learning for VisionKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Graph Algorithms and Modeling for the Web: Graph embeddings and representation learning for Web-related graphsSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsWeb Mining and Content Analysis: Bridging structured and unstructured data
📝 Abstract
Embeddings serve as condensed vector representations for real-world entities, finding applications in Natural Language Processing (NLP), Computer Vision, and Data Management across diverse downstream tasks. Here, we introduce novel specialized embeddings optimized, and explicitly tailored to encode the intricacies of complex 2-D context in tables, featuring horizontal, vertical hierarchical metadata, and nesting. To accomplish that we define the Bi-dimensional tabular coordinates, separate horizontal, vertical metadata and data contexts by introducing a new visibility matrix, encode units and nesting through the embeddings specifically optimized for mimicking intricacies of such complex structured data. Through evaluation on 5 large-scale structured datasets and 3 popular downstream tasks, we observed that our solution outperforms the state-of-the-art models with the significant MAP delta of up to 0.28. GPT-4 LLM+RAG slightly outperforms us with MRR delta of up to 0.1, while we outperform it with the MAP delta of up to 0.42.
Problem

Research questions and friction points this paper is trying to address.

Optimize embeddings for 2-D table structures.
Encode horizontal and vertical hierarchical metadata.
Improve performance on structured dataset tasks.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Specialized embeddings for 2-D tabular data
Visibility matrix for metadata separation
Evaluation on large-scale structured datasets
🔎 Similar Papers
No similar papers found.
G
Gyanendra Shrestha
Florida State University, Tallahassee, Florida, USA
Chutian Jiang
Chutian Jiang
Florida State University, Tallahassee, Florida, USA
S
Sai Akula
Florida State University, Tallahassee, Florida, USA
V
Vivek Yannam
Florida State University, Tallahassee, Florida, USA
A
A. Pyayt
University of South Florida, Tampa, Florida, USA
M
M. Gubanov
Florida State University, Tallahassee, Florida, USA