Evaluating Multi-Dimensional Generalization of Large Language Models in Temporal Extraction Tasks

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient generalization caused by domain shift and the limitations of one-sided evaluation in temporal extraction tasks using large language models. We propose a four-dimensional distribution shift generalization framework that systematically compares the cross-domain transfer performance of multiple model configurations under inductive, deductive, and abductive prompting strategies. Our findings reveal that the correlation between base performance and generalization capability diminishes as distributional shifts intensify, while inductive prompting demonstrates significant robustness advantages under complex distribution shifts. By overcoming the constraints of single-dimensional evaluation, this work provides both theoretical foundations and practical guidance for enhancing the reliable generalization of large language models in temporal extraction.
📝 Abstract
Time and event expression extraction are fundamental temporal reasoning tasks, but the problem remains difficult due to annotation ambiguity, domain sensitivity, and unstable model behavior. Existing evaluations focus on in-domain performance, offering limited insight into reliability under distribution shifts. We evaluate multiple model configurations across families, architectures, and reasoning strategies over four dimensions of generalization, examining transfer from base performance, cross-dimensional correlations, and the effects of scale, architecture, and prompting. This provides a systematic study of how prompted LLMs generalize in time and event expression extraction tasks. We find that strong base-task performance generally predicts better generalization. However, this relationship weakens under substantial distribution shifts. Inductive prompting performs most consistently across domain shift, adversarial perturbations, compositionality, and length increase, while gains from scale, architecture, and deductive and abductive prompting strategies are uneven and dimension-specific. We conclude that LLM generalization in temporal extraction tasks cannot be predicted from any single dimension alone and cannot be reliably inferred from in-domain or single-dimension evaluations, highlighting the need for reasoning strategies that generalize across dimensions.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Temporal Extraction
Multi-Dimensional Generalization
Time and Event Expression Extraction
Distribution Shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Temporal Extraction
Multi-Dimensional Generalization
Inductive Prompting
Distribution Shifts