Randomness in large language models: What researchers need to know (and report)

📅 2026-07-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the threat posed by the stochasticity of large language model (LLM) outputs to research reproducibility, demonstrating that randomness arises not only from sampling strategies but also from non-sampling factors such as silent model updates, numerical rounding, and expert routing in mixture-of-experts architectures. The work is the first to explicitly model LLM outputs as draws from a probability distribution and systematically evaluates the impact of various sources of randomness on downstream outcomes through regression analysis in a sentiment classification task. Empirical comparisons across multiple software and hardware environments—using both API-accessible and locally deployed open-source models—reveal substantial output variability even at zero temperature. Building on these findings, the paper proposes standardized reporting guidelines for research papers and reproduction packages to encourage the adoption of stricter LLM usage protocols within the community.
📝 Abstract
Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.
Problem

Research questions and friction points this paper is trying to address.

randomness
large language models
reproducibility
LLM outputs
research reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

randomness
reproducibility
large language models
reporting standards
stochastic outputs
💼 Related Jobs
No related jobs found.
G
Guillaume Coqueret
EMLYON Business School, France
Joan Llull
Joan Llull
IAE-CSIC and Barcelona School of Economics
EconomicsLabor economicsStructural MicroeconometricsEconomics of MigrationApplied
F
Florian Oswald
University of Turin and Collegio Carlo Alberto, Italy
C
Christophe Pérignon
HEC Paris, France
C
Christoph Scheuch
Humboldt University of Berlin, Germany
Lars Vilhuber
Lars Vilhuber
Cornell University
Labor Economics