Measuring Iterative Temporal Reasoning with Time Puzzles

πŸ“… 2026-01-12
πŸ›οΈ arXiv.org
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the capacity of large language models to perform iterative temporal reasoning without external tools, revealing significant limitations in handling cross-cultural calendar constraints and dynamic temporal anchors. To this end, the authors introduce Time Puzzles, the first benchmark for constrained date reasoning that supports multiple valid solutions, integrates cross-cultural calendar rules, and enables dynamic puzzle generation through constraint satisfaction problem modeling and algorithmic puzzle synthesis. Experiments across 13 prominent models demonstrate that even GPT-5 achieves only 49.3% accuracy without tool assistance, while all other models score below 31%. Performance improves markedly when explicit date rewriting or web search is incorporated, underscoring a critical gap in current models’ ability to reliably invoke appropriate reasoning tools.

Technology Category

Knowledge Representation and Reasoning: Computational Complexity of ReasoningCognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
πŸ“ Abstract
We introduce Time Puzzles, a constraint-based date inference task for evaluating iterative temporal reasoning. Each puzzle combines factual temporal anchors with (cross-cultural) calendar relations, admits one or multiple valid solution dates, and is algorithmically generated for controlled, dynamic, and continual evaluation. Across 13 diverse LLMs, Time Puzzles well distinguishes their iterative temporal reasoning capabilities and remains challenging without tools: GPT-5 reaches only 49.3% accuracy and all other models stay below 31%, despite the dataset's simplicity. Web search consistently yields substantial gains and using code interpreter shows mixed effects, but all models perform much better when constraints are rewritten with explicit dates, revealing a gap in reliable tool use. Overall, Time Puzzles presents a simple, cost-effective diagnostic for tool-augmented iterative temporal reasoning.
Problem

Research questions and friction points this paper is trying to address.

iterative temporal reasoning
date inference
Time Puzzles
large language models
constraint-based reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

iterative temporal reasoning
Time Puzzles
constraint-based date inference
tool-augmented reasoning
algorithmic evaluation
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
Z
Zhengxiang Wang
Department of Linguistics & IACS, Stony Brook University
Zeyu Dong
Zeyu Dong
boston university
Computer VisionNLPDeep Learning