🤖 AI Summary
This study addresses the bottleneck of dynamically localizing rows, columns, and cells in table reasoning with multimodal large language models by proposing a structure-aware framework. The framework interleaves chain-of-thought generation with bounding-box cropping tool invocations to precisely align reasoning steps with corresponding visual regions. Specifically, this work constructs the InterTab-22K dataset, designs Structure-Sensitive Alignment (SSA) and Area-Level Optimization (ALO) strategies, and adopts a two-stage training paradigm to enhance model capabilities. Experimental results demonstrate that the proposed method improves average accuracy from 68.28% to 73.17% across nine benchmarks, significantly outperforming existing approaches.
📝 Abstract
Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-, column-, and cell-level evidence as the question unfolds. Encoder-side table structure and generic interleaved visual chain-of-thought still do not bind each reasoning step to that structure. We propose \textbf{InterTab}, an \textbf{Inter}leaved structure-aware framework for CoT reasoning over \textbf{Tab}le images, interleaves chain-of-thought with tool calls that crop structure-aligned table regions. First, we build InterTab-22K, includes reasoning trajectories in which each step is tied to both a structural location and a bounding box. InterTab is trained in two stages: supervised structure-aware alignment (SSA) on InterTab-22K teaches the model to interleave reasoning with structure-aligned crops, and active localization optimization (ALO) further optimizes answer correctness, localization IoU, and output format, while penalizing missing or excessive tool calls. Experiments on nine table benchmarks show that InterTab improves the average accuracy of its backbone from 68.28% to 73.17% and achieves the best average performance among all compared methods. Code and data will be released soon.