🤖 AI Summary
This study addresses the significant non-deterministic bias introduced by the implicit injection of the current date within system prompts during large language model (LLM) evaluation, which severely compromises reproducibility and fairness. Through multi-task evaluations across nine LLMs and six datasets, combined with ablation studies and statistical tests, this work identifies the hidden date as a primary source of non-determinism and demonstrates that standard techniques such as chain-of-thought prompting fail to mitigate—and may even amplify—this effect. The findings quantify performance fluctuations of up to 14% in question answering, mathematics, code generation, and translation tasks, revealing that date-induced variance substantially exceeds traditional noise sources like batch size and leads to model ranking drift. These insights provide critical empirical evidence for refining standardized evaluation protocols.
📝 Abstract
Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.