Score
Designs, builds, and evaluates systems that perform prediction or classification on tabular data by providing example rows at inference time (in‑context) instead of fine‑tuning model weights. This work includes engineering encoders and prompt formats to represent table features, selecting and ordering in‑context examples, and integrating or adapting frozen backbone models and lightweight adapters to enable rapid adaptation and low‑latency inference on new tabular tasks.
Tabular data exhibit strong heterogeneity, and existing models suffer from poor generalization and difficulty in zero-shot adaptation to new tasks. Method: We propose the Discriminative Tabular Pre-trained Transformer (TabDPT), the first framework integrating real-table-driven self-supervised pretraining with retrieval-augmented in-context learning (ICL). It introduces numerical-aware embedding and attention mechanisms, alongside a lightweight discriminative architecture. Contribution/Results: TabDPT achieves true zero-shot cross-task generalization without fine-tuning—overcoming a key bottleneck in large language models’ handling of structured numerical tables. It attains state-of-the-art zero-shot performance on the CC18 classification and CTR23 regression benchmarks. Performance scales consistently with both model and data size, while maintaining efficient inference and strong scalability.
Existing approaches to modeling medical tabular data often neglect column-level contextual information (e.g., header semantics), require labor-intensive, manual preprocessing, and lack semantic interpretability. Method: We propose TabText—a novel framework that systematically encodes tabular structure into promptable natural language text, enabling end-to-end semantic representation learning via large language models (LLMs). TabText integrates structure-aware prompt engineering with a multi-task health prediction fine-tuning paradigm, supporting zero- or light-preprocessing modeling while remaining compatible with conventional feature fusion. Results: Evaluated on nine clinical prediction tasks, TabText alone establishes high-performance, lightweight baselines. When fused with traditional features, it yields an average AUC improvement of 6%, with even the worst-case gain reaching 6%, significantly enhancing model robustness and generalization across diverse healthcare scenarios.
This work investigates the transferability of task-agnostic embeddings for tabular data, systematically evaluating generative foundation models (e.g., TabPFN, TabICL) against traditional feature engineering—specifically TableVectorizer—on downstream tasks including anomaly detection and supervised learning. The study finds that TableVectorizer, leveraging lightweight, hand-crafted feature transformations, produces high-quality universal embeddings that match or exceed the performance of large foundation models across multiple benchmarks, while achieving inference speeds three orders of magnitude faster. These results challenge the prevailing assumption that complex foundation models are inherently superior for generic tabular representation learning. Crucially, this is the first empirical demonstration that simple, interpretable, and computationally efficient feature engineering can be highly competitive in learning task-agnostic tabular embeddings. The work thus establishes a more practical, deployable paradigm for tabular representation learning—grounded in transparency, efficiency, and robust generalization.
Existing large language models (LLMs) exhibit limited performance on tabular prediction tasks—including classification, regression, and missing-value imputation—primarily due to their lack of native modeling capacity for structured semantics and numerical relationships. Method: We propose the first large-scale instruction-tuning paradigm tailored for tabular understanding, built upon Llama-2. It leverages a comprehensively curated, multi-scenario instruction-based tabular corpus and systematically applies supervised fine-tuning and in-context learning adaptation. Contribution/Results: We demonstrate, for the first time, that a single LLM can robustly and uniformly address all three core tabular prediction tasks under zero-shot, few-shot, and in-context learning settings. Our approach achieves significant improvements over traditional models (e.g., XGBoost, TabPFN) and prior LLM-based methods across 12 benchmarks—averaging +8.2% accuracy, +6.5% R², and −12.4% imputation MAE—establishing a new state-of-the-art for LLM-driven structured data intelligence.
This work addresses the challenge of simultaneously ensuring fidelity, security, and utility in tabular data generation. We systematically evaluate five mainstream generative models—diffusion models, GANs, VAEs, CTGAN, and TVAE—across 16 real-world datasets. To reduce tuning overhead while preserving near-optimal performance, we propose a model-specific, streamlined hyperparameter search space. Our analysis reveals that diffusion models’ advantages are highly sensitive to tuning intensity and vanish under constrained GPU budgets. We further introduce the first joint optimization of multiple encoding strategies (e.g., embedding and one-hot) with model architectures. Finally, we establish a standardized, reproducible benchmark framework: all models undergo dataset-specific hyperparameter tuning, yielding consistent improvements in generation quality; evaluation results and full tuning protocols are publicly released.
为解决表格数据预测中准确性和效率的平衡问题,提出Tydra模型,结合Transformer和状态空间模型,减少推理时间同时保持较高预测性能。
This work addresses a critical oversight in modern tabular learning benchmarks—the neglect of feature engineering—which introduces bias in model evaluation and obscures its pivotal role in real-world applications. To bridge this gap, the authors propose TabPrep, a lightweight, domain-knowledge-driven preprocessing pipeline that systematically enhances model performance by identifying and transforming three fundamental structural data patterns through targeted feature generators. TabPrep integrates systematic feature engineering into mainstream tabular benchmarking for the first time, demonstrating that well-designed preprocessing alone can surpass performance gains from most architectural innovations. Evaluated on the TabArena benchmark, TabPrep consistently boosts the accuracy of diverse models—including tree-based methods, neural networks, linear models, and foundation models—while maintaining low computational overhead and strong generalizability.
This work addresses the trade-off between predictive performance and inference efficiency in existing tabular prediction models, which often sacrifice speed for accuracy, hindering deployment in resource-constrained or latency-sensitive settings. The authors propose an efficient foundation model for tabular data that eschews retrieval mechanisms and instead introduces row-wise attention, combined with long-context pretraining, architectural optimizations, and self-supervised learning on large-scale real-world tabular datasets. The resulting model achieves prediction performance comparable to TabDPT v1.1 on the TabArena-Lite, CC18, and CTR23 benchmarks while accelerating inference by several orders of magnitude. This approach strikingly balances effectiveness and efficiency, establishing a new state-of-the-art as the fastest general-purpose tabular prediction model to date.
Existing tabular in-context learning methods often couple feature representations with specific prediction targets, limiting their generalization across diverse tasks. This work proposes a task-agnostic encoder-decoder architecture that employs a single shared encoder to learn universal row embeddings from unlabeled real-world tables, paired with multiple task-specific decoders to support six downstream tasks: classification, regression, anomaly detection, clustering, entity matching, and entity classification in relational databases. To the best of our knowledge, this is the first approach to enable effective transfer of a unified representation across multiple tabular in-context learning tasks. The method achieves state-of-the-art performance on most tasks and substantially enhances model generalization.
本文提出ARASH方法,通过局部邻域分析选择最优样本,提高表格预测效率,减少TabPFN的提示长度和内存使用。