TabText: A Flexible and Contextual Approach to Tabular Data Representation

πŸ“… 2022-06-21
πŸ“ˆ Citations: 8
✨ Influential: 1
πŸ“„ PDF

career value

159K/year
πŸ€– AI Summary
Existing approaches to modeling medical tabular data often neglect column-level contextual information (e.g., header semantics), require labor-intensive, manual preprocessing, and lack semantic interpretability. Method: We propose TabTextβ€”a novel framework that systematically encodes tabular structure into promptable natural language text, enabling end-to-end semantic representation learning via large language models (LLMs). TabText integrates structure-aware prompt engineering with a multi-task health prediction fine-tuning paradigm, supporting zero- or light-preprocessing modeling while remaining compatible with conventional feature fusion. Results: Evaluated on nine clinical prediction tasks, TabText alone establishes high-performance, lightweight baselines. When fused with traditional features, it yields an average AUC improvement of 6%, with even the worst-case gain reaching 6%, significantly enhancing model robustness and generalization across diverse healthcare scenarios.
πŸ“ Abstract
Tabular data is essential for applying machine learning tasks across various industries. However, traditional data processing methods do not fully utilize all the information available in the tables, ignoring important contextual information such as column header descriptions. In addition, pre-processing data into a tabular format can remain a labor-intensive bottleneck in model development. This work introduces TabText, a processing and feature extraction framework that extracts contextual information from tabular data structures. TabText addresses processing difficulties by converting the content into language and utilizing pre-trained large language models (LLMs). We evaluate our framework on nine healthcare prediction tasks ranging from patient discharge, ICU admission, and mortality. We show that 1) applying our TabText framework enables the generation of high-performing and simple machine learning baseline models with minimal data pre-processing, and 2) augmenting pre-processed tabular data with TabText representations improves the average and worst-case AUC performance of standard machine learning models by as much as 6%.
Problem

Research questions and friction points this paper is trying to address.

Converting tabular medical data into contextual language representations
Leveraging LLMs to generate task-independent embeddings for predictions
Improving predictive accuracy for healthcare tasks using contextual information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Converts tabular data into contextual language representations
Uses pretrained language models for task-independent embeddings
Enhances predictive accuracy for challenging healthcare tasks
πŸ”Ž Similar Papers
No similar papers found.
K
Kimberly Villalobos Carballo
Tandon School of Engineering, New York University
L
Liangyuan Na
Operations Research Center, Massachusetts Institute of Technology
Yu Ma
Yu Ma
Indiana University
Computer Science
L
L. Boussioux
Foster School of Business, University of Washington, Foster School of Business
C
C. Zeng
Stern School of Business, New York University Abu Dhabi
L
L. Soenksen
Abdul Latif Jameel Clinic for Machine Learning in Health, Massachusetts Institute of Technology
D
D. Bertsimas
Sloan School of Management, Massachusetts Institute of Technology