Understanding Errors in LLM-Based Question Answering over Imperfect Tables

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the vulnerability of large language models to row-order bias in error-containing table question answering, where merely localizing errors proves insufficient for ensuring accuracy. To tackle this, we propose GBDI, a workflow that introduces a novel geometrically balanced discovery and intervention strategy. By integrating multi-view ensemble reasoning, table reordering, and code execution, this method systematically reveals and mitigates the bias imposed by row order on error discovery. Experimental results demonstrate that GBDI improves accuracy by 3.8 to 18.5 percentage points over baselines on the RADAR-T benchmark. This work establishes a new paradigm for robust table question answering in the presence of tabular errors.
πŸ“ Abstract
We investigate error discovery and handling in question answering over imperfect tables through controlled studies across three large language models (LLMs) on human-reviewed RADAR-T examples. Answering questions over these tables requires handling errors that can affect the answer. We vary row order and compare original, error-marked, and repaired tables to test whether discovery depends on where errors appear and whether providing their locations is sufficient for accurate question answering. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. Complete discovery is higher for back than front placements and, averaged over the tested mean positions, for compact than widely spaced layouts. Second, providing verified error locations alone is insufficient for accurate QA, leaving a substantial accuracy gap between error-marked and repaired tables. Providing tables with human-reviewed repairs already applied raises code-assisted QA accuracy by 39.0-59.1 percentage points over the error-marked tables across the three systems. GBDI, a simple workflow, puts these findings into practice by combining error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors. On RADAR-T, GBDI raises observed QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five systems. These results highlight the importance of both reliable error discovery and effective error handling in question answering over imperfect tables. Our anonymous repository is available at https://anonymous.4open.science/r/GBDI-ICLR-2027-85BD/
Problem

Research questions and friction points this paper is trying to address.

imperfect tables
question answering
large language models
error discovery
error handling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Imperfect Tables
Error Discovery
Question Answering
Large Language Models
GBDI
πŸ”Ž Similar Papers
No similar papers found.