🤖 AI Summary
This study addresses the challenges of inaccurate workbook localization and incomplete computation in question answering over spreadsheet collections by proposing FiCo, a method that innovatively integrates semantic source selection with schema-based precise calculation. Specifically, FiCo accurately identifies data sources through document summary retrieval and similar workbook disambiguation, subsequently executing constrained SQL queries over the entire selected tables to perform computations. This design effectively overcomes performance bottlenecks caused by erroneous source selection. Experimental evaluations demonstrate that FiCo achieves an accuracy of 76.2% on the DataBench dataset, representing a 9.9% improvement over baseline methods, and attains 79.7% on the MiMoTable dataset. These results indicate that the proposed approach significantly outperforms existing prefix-based Retrieval-Augmented Generation (RAG) methods.
📝 Abstract
Question answering over spreadsheet collections requires finding the correct workbook and computing over complete tables. We introduce Find-then-Compute (FiCo), which retrieves document summaries, disambiguates similar workbooks, and executes constrained Structured Query Language (SQL) over the selected full table. On DataBench (80 datasets, 1,810 questions), FiCo reaches 76.2% accuracy: 9.9 points above a strong TableRAG-style baseline on the same frozen workbook choices (66.3%) under the tracks'prespecified evaluators, and 63.4 points above prefix RAG (12.8%). On 508 MiMoTable questions, FiCo reaches 79.7%, versus 22.2% for prefix RAG. Giving the strong baseline the gold workbook raises it from 66.3% to 76.3%, exposing a 10.0-point source-selection cost under fixed compute. Despite 95.1% document recall and 98.5% executable SQL, only 81.3% of questions execute on the gold workbook. FiCo's advantage comes from integrating semantic source selection with exact, schema-grounded computation.