Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
Reproducibility is essential for scientific research, but prior work has shown that LLM outputs can vary with hardware and batching. This paper identifies another overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day.
Across 9 recent LLMs and 6 datasets spanning multiple-choice QA, math reasoning, code generation, and machine translation, performance varied solely with the current date. The reported deltas reached up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation, and model rankings also shifted, affecting leaderboards.
The date effect exceeded other sources of non-determinism such as batch size and numerical precision. Standard prompting techniques, including chain-of-thought and few-shot prompting, did not reduce the sensitivity, and chain-of-thought even amplified it. The authors conclude that careful evaluation protocols are needed to ensure reproducibility and fair comparisons in LLM research.