Evidence for Daily and Weekly Periodic Variability in GPT-4o Performance
This study investigates whether the performance of large language models remains constant over time by conducting a longitudinal replication experiment on GPT-4o over three months, with ten evaluations every three hours under identical physics tasks. Employing time-series data collection, controlled experimental design, and Fourier spectral analysis, the research reveals—for the first time—significant diurnal and weekly periodic fluctuations in model performance, with approximately 20% of output variance attributable to these rhythmic patterns. These findings challenge the foundational assumption of temporal invariance in model performance, demonstrating that the output quality of large language models exhibits systematic time dependence. The results carry important implications for the reproducibility of AI research and the prevailing paradigms used to evaluate model capabilities.