Evidence for Daily and Weekly Periodic Variability in GPT-4o Performance

📅 2026-02-06
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the performance of large language models remains constant over time by conducting a longitudinal replication experiment on GPT-4o over three months, with ten evaluations every three hours under identical physics tasks. Employing time-series data collection, controlled experimental design, and Fourier spectral analysis, the research reveals—for the first time—significant diurnal and weekly periodic fluctuations in model performance, with approximately 20% of output variance attributable to these rhythmic patterns. These findings challenge the foundational assumption of temporal invariance in model performance, demonstrating that the output quality of large language models exhibits systematic time dependence. The results carry important implications for the reproducibility of AI research and the prevailing paradigms used to evaluate model capabilities.

Technology Category

Natural Language Processing: (Large) Language ModelsMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Computational Creativity

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Large language models for searchGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language models (LLMs) are increasingly used in research both as tools and as objects of investigation. Much of this work implicitly assumes that LLM performance under fixed conditions (identical model snapshot, hyperparameters, and prompt) is time-invariant. If average output quality changes systematically over time, this assumption is violated, threatening the reliability, validity, and reproducibility of findings. To empirically examine this assumption, we conducted a longitudinal study on the temporal variability of GPT-4o's average performance. Using a fixed model snapshot, fixed hyperparameters, and identical prompting, GPT-4o was queried via the API to solve the same multiple-choice physics task every three hours for approximately three months. Ten independent responses were generated at each time point and their scores were averaged. Spectral (Fourier) analysis of the resulting time series revealed notable periodic variability in average model performance, accounting for approximately 20% of the total variance. In particular, the observed periodic patterns are well explained by the interaction of a daily and a weekly rhythm. These findings indicate that, even under controlled conditions, LLM performance may vary periodically over time, calling into question the assumption of time invariance. Implications for ensuring validity and replicability of research that uses or investigates LLMs are discussed.
Problem

Research questions and friction points this paper is trying to address.

time invariance
large language models
reproducibility
performance variability
longitudinal study
Innovation

Methods, ideas, or system contributions that make the work stand out.

time invariance
periodicity
large language models
Fourier analysis
reproducibility
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Paul Tschisgale
Leibniz Institute for Science and Mathematics Education, Kiel, Germany
P
Peter Wulff
Ludwigsburg University of Education, Ludwigsburg, Germany