Investigating Human--AI Discrepancies via Multiple-Solution Problems

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the divergence in reasoning distributions between humans and AI on problems admitting multiple valid solutions, aiming to transcend the limitations of single-metric accuracy evaluation. Methodologically, we construct a benchmark of multi-solution reasoning puzzles and propose an evaluation framework grounded in probability distributions rather than binary correctness. By integrating statistical distribution analysis with prompt engineering perturbation experiments, we systematically compare the solution spaces of human participants and frontier AI models. Our findings reveal that AI models exhibit highly convergent distributions with significantly lower cognitive diversity than humans. This work uncovers fundamental discrepancies in cognitive preferences between humans and machines, highlights the inadequacy of conventional evaluation metrics, and offers a novel perspective for the safe deployment of AI systems.
📝 Abstract
Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and model reasoning lead to different distributions over valid solutions. Our testbed comprises 270 reasoning puzzles across five puzzle families. These multiple-solution puzzles each have 3 to 8 valid solutions and are simple enough that humans and models can solve them reliably. The resulting distributions differ markedly: models differ from one another, yet resemble each other far more than they resemble humans. Model distributions are, moreover, within every puzzle family, less diverse than human ones. We compare these discrepancies across puzzle categories, and trace how they respond to reasoning-effort settings, to prompting, and to perturbations of the puzzle that leave its solutions unchanged. Together, these results point at significant differences between human and AI problem-solving processes, and their choice among equally defensible solutions. As progressive deployment of AI systems in society comes into focus, evaluating such differences (beyond one-dimensional accuracy metrics) is increasingly important. Data and code are available at https://hai-discrepancies.github.io/
Problem

Research questions and friction points this paper is trying to address.

Human-AI discrepancies
Multiple-solution problems
Reasoning
Solution distribution
AI evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-AI Discrepancies
Multiple-Solution Problems
Solution Distribution Diversity
Reasoning Evaluation
Prompting and Perturbation Analysis
💼 Related Jobs
No related jobs found.
Z
Zihao Wang
Department of Mathematics, Stanford University
F
Francesco Insulla
Institute of Computational and Mathematical Engineering, Stanford University
Andrea Montanari
Andrea Montanari
John D. and Sigrid Banks Professor, Statistics and Mathematics, Stanford University
statisticsmachine learningprobability theoryinformation theorysignal processing