Overwhelmed by Choice: Studying LLM Decision Making at Scale

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study reveals the performance degradation and evaluation bias of large language models (LLMs) in decision-making tasks involving large-scale candidate sets. By establishing candidate set size as a critical evaluation variable, we demonstrate that performance advantages observed in small-scale settings do not guarantee robustness at scale, identifying confidence collapse and positional bias as two primary failure modes. To address these challenges, this work proposes a hierarchical partitioning strategy and a permutation-based reasoning method. Experimental results indicate that these approaches effectively mitigate performance degradation, improving accuracy by approximately 20 percentage points when the number of candidates reaches N=160. Ultimately, this research establishes a new paradigm for the reliable application and rigorous evaluation of LLMs in large-scale decision-making scenarios.
📝 Abstract
Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically evaluate LLMs as the number of competing candidates increases and find substantial accuracy degradation across tasks, prompting strategies, and model scales. Controlled analyses show that standard long-context retrieval explanations cannot fully account for this degradation. Instead, we identify two systematic failure patterns. First, gold-margin collapse: the score gap between the correct answer and the strongest distractor progressively shrinks, driven primarily by weakening confidence in the correct answer. Second, earlier candidate preferences become increasingly difficult to overturn, with later candidates exerting progressively weaker influence on the final prediction. Motivated by these findings, we evaluate hierarchical partitioning and permutation-based inference, which improve accuracy by roughly 20 percentage points at $N=160$ on both HotpotQA and MIMIC. Overall, our results identify candidate-set scale as an important evaluation-protocol variable and show that strong small-option performance does not necessarily imply robust large-scale candidate comparison.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Decision Making
Candidate Selection
Scaling
Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Candidate-set Scaling
Decision Making
Hierarchical Partitioning
Permutation-based Inference
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yu-Chi Lin
University of California, Los Angeles
Aryan Seth
Aryan Seth
Birla Institute of Technology and Science, Pilani
A
Anshul Aravind
University of California, Los Angeles
Eugene Lee
Eugene Lee
University of Cincinnati
Machine Learning
Tanmay Parekh
Tanmay Parekh
Student, University of California Los Angeles
Natural Language ProcessingMachine Learning
N
Nanyun Peng
University of California, Los Angeles
K
Kai-Wei Chang
University of California, Los Angeles