🤖 AI Summary
This study addresses the absence of efficiency optimization benchmarks for the two-dimensional attention mechanisms in tabular foundation models such as TabPFN. We present the first systematic exploration of efficient implementations for table-native row-column attention. Reproducible benchmarks are established on A100, H100, and B200 GPUs to evaluate the forward and backward propagation throughput of backends including FlashAttention, Torch SDPA, and vLLM under realistic data shapes. Our analysis reveals that the optimal backend varies dynamically with hardware architecture and sequence length: FlashAttention generally achieves superior overall performance, whereas CuDNN demonstrates advantages in specific column-attention scenarios. The code has been released as open source, establishing a foundational framework for accelerating the inference and training of future tabular models.
📝 Abstract
Tabular in-context learners such as TabPFN, Mitra, or ConTextTab rely on alternating row and column attention over 2D sequences of latent embeddings. These attention patterns differ markedly from the one-dimensional case in language models: row attention involves longer sequences while column attention operates on much shorter ones, and the strided memory layout of tabular data makes producing contiguous tensors costly. Moreover, the hidden dimensions used in current models are small compared to recent language models. Yet efficient attention has been studied mostly for one-dimensional sequences, leaving the two-dimensional tabular setting unexplored. To this end, we create a reproducible benchmarking setup and study the unique characteristics of tabular attention across several backends -- Torch SDPA (efficient and cuDNN), FlashAttention-2/3/4, and the inference-only backends vLLM and SageAttention -- measuring forward and backward throughput across realistic tabular shapes on three GPU generations (A100, H100, B200). We find that the optimal backend choice differs between column and row attention and varies across hardware as well as model specifics: While the FlashAttention implementations tailored for each GPU generation perform overall best, they are at times outperformed by CuDNN in the case of column attention at longer sequences with cross-over points depending on the head dimension. Among inference-only backends, SageAttention performs well for row attention and large sequences beyond 16\,k rows. Our reproducible benchmark lays the foundation for future improvements to table-native attention. The self-contained benchmarking and evaluation code is openly available at: https://github.com/SAP-samples/tabular-attention-benchmark