🤖 AI Summary
This study investigates whether existing temporal question answering (TSQA) systems genuinely leverage numerical evidence for reliable reasoning. To address this, we propose a unified evaluation framework and construct the COMMON-TSQA benchmark, incorporating six categories of intervention experiments alongside a rationale auditing mechanism to rigorously assess systems across input sensitivity, factual grounding, and logical consistency. Our findings reveal that high accuracy scores frequently mask underlying deficiencies in evidence utilization and fallback behaviors, while generated rationales often contain unsupported numerical claims. By exposing these discrepancies between surface-level performance and actual reasoning fidelity, this work provides both a rigorous methodological approach and a new benchmark for diagnosing the genuine reasoning capabilities of TSQA systems.
📝 Abstract
In recent years, time-series question answering (QA) systems have made significant progress. However, generating a correct answer does not show whether retaining the supplied numerical series improves task performance, nor whether the prediction is sensitive to changes in that input. While some systems provide rationales, answer accuracy also does not show whether their numerical claims are grounded in the supplied series or whether the stated inference is valid. In this work, we focus on evaluating four time-series QA systems: TimeOmni-1, ChatTS, TimeOmni-VL, and Time-MQA. First, for three systems with released evaluation data, we reproduce their reported results and compare the performance of the systems with their backbones. Then, we introduce a benchmark named COMMON-TSQA, which collects public evaluation datasets from existing time-series benchmarks and unifies their sample representation, task definitions, and answer schemas, while evaluating each system through its own interface under common evaluation criteria. The evaluation uses the original condition and six interventions while keeping the question and target fixed. Our analysis shows that aggregate performance alone can obscure how systems use numerical evidence. Similar task-level scores can arise despite substantial changes in individual predictions. Some interventions induce simple fallback behavior rather than preserved task ability. We also evaluate rationales for factual grounding, inference validity, and consistency with the final answer. We find that rationales often contain time-series claims unsupported by the input. Moreover, the rationale audit shows that agreement between a rationale and its final answer can coexist with incorrect numerical descriptions or invalid intermediate inferences.