Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant non-deterministic bias introduced by the implicit injection of the current date within system prompts during large language model (LLM) evaluation, which severely compromises reproducibility and fairness. Through multi-task evaluations across nine LLMs and six datasets, combined with ablation studies and statistical tests, this work identifies the hidden date as a primary source of non-determinism and demonstrates that standard techniques such as chain-of-thought prompting fail to mitigate—and may even amplify—this effect. The findings quantify performance fluctuations of up to 14% in question answering, mathematics, code generation, and translation tasks, revealing that date-induced variance substantially exceeds traditional noise sources like batch size and leads to model ranking drift. These insights provide critical empirical evidence for refining standardized evaluation protocols.
📝 Abstract
Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift, affecting leaderboards. This date effect exceeds other sources of non-determinism, such as batch size and numerical precision. Standard prompting techniques -- chain-of-thought and few-shot prompting -- do not reduce the sensitivity; chain-of-thought even amplifies it. Our findings underscore the need for careful evaluation protocols to ensure reproducibility and fair comparisons in LLM research.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Reproducibility
System Prompts
LLM Evaluation
Non-determinism
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Reproducibility
System Prompts
Evaluation Protocol
Non-determinism