Large Language Model Strategic Reasoning Evaluation through Behavioral Game Theory

📅 2025-02-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenge of evaluating strategic reasoning capabilities in large language models (LLMs), where contextual confounds hinder accurate assessment. To this end, it introduces the first evaluation framework grounded in behavioral game theory—designed to isolate and precisely measure strategic reasoning by controlling for contextual effects. Methodologically, the framework integrates multi-round strategic games, chain-of-thought prompt attribution analysis, and demographic feature encoding to systematically examine the impacts of model scale, prompt engineering, and latent sociodemographic attributes. Key contributions include: (1) pioneering the application of behavioral game theory to LLM reasoning evaluation; (2) uncovering a non-monotonic relationship between model size and strategic performance; and (3) revealing systematic reasoning biases correlated with gender, sexual orientation, and other implicit social attributes—providing empirical evidence of latent bias. Validated across 22 state-of-the-art models, the framework identifies GPT-4o and DeepSeek-R1 as top performers, establishing a novel paradigm for fairness-aware and interpretable LLM evaluation.

Technology Category

Game Theory and Economic Paradigms: Behavioral Game TheoryMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language Models

Application Category

Economics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactionsSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
Strategic decision-making involves interactive reasoning where agents adapt their choices in response to others, yet existing evaluations of large language models (LLMs) often emphasize Nash Equilibrium (NE) approximation, overlooking the mechanisms driving their strategic choices. To bridge this gap, we introduce an evaluation framework grounded in behavioral game theory, disentangling reasoning capability from contextual effects. Testing 22 state-of-the-art LLMs, we find that GPT-o3-mini, GPT-o1, and DeepSeek-R1 dominate most games yet also demonstrate that the model scale alone does not determine performance. In terms of prompting enhancement, Chain-of-Thought (CoT) prompting is not universally effective, as it increases strategic reasoning only for models at certain levels while providing limited gains elsewhere. Additionally, we investigate the impact of encoded demographic features on the models, observing that certain assignments impact the decision-making pattern. For instance, GPT-4o shows stronger strategic reasoning with female traits than males, while Gemma assigns higher reasoning levels to heterosexual identities compared to other sexual orientations, indicating inherent biases. These findings underscore the need for ethical standards and contextual alignment to balance improved reasoning with fairness.
Problem

Research questions and friction points this paper is trying to address.

Evaluate LLMs' strategic reasoning using behavioral game theory.
Assess impact of model scale and prompting on strategic decision-making.
Investigate demographic biases in LLMs' strategic reasoning patterns.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavioral game theory evaluates LLM strategic reasoning.
Chain-of-Thought prompting enhances selective model performance.
Demographic features influence LLM decision-making patterns.
🔎 Similar Papers
No similar papers found.