Prompting Science Report 2: The Decreasing Value of Chain of Thought in Prompting

📅 2025-06-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
The widespread assumption of chain-of-thought (CoT) prompting’s universal effectiveness lacks rigorous empirical validation across diverse models and tasks. Method: We conduct a systematic, multi-model, multi-task evaluation using standardized benchmarks, token-level cost analysis, error-pattern statistics, and controlled cross-model experiments. Contribution/Results: We find—contrary to prevailing assumptions—that CoT yields negligible performance gains for native reasoning models and only marginal improvements (≤1.2% accuracy) for non-reasoning models, at the cost of reduced accuracy stability. It incurs an average 47% increase in response latency and a 3.8× rise in token consumption, while introducing additional logical errors. Our work is the first to empirically refute CoT’s general efficacy, identifying three critical practical bottlenecks: diminishing returns on performance gain, increased output volatility, and sharply escalated inference overhead—providing essential evidence for rational prompt engineering.

Technology Category

Knowledge Representation and Reasoning: Computational Complexity of ReasoningCognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningNatural Language Processing: Prompt Engineering / Prompting

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
This is the second in a series of short reports that seek to help business, education, and policy leaders understand the technical details of working with AI through rigorous testing. In this report, we investigate Chain-of-Thought (CoT) prompting, a technique that encourages a large language model (LLM) to"think step by step"(Wei et al., 2022). CoT is a widely adopted method for improving reasoning tasks, however, our findings reveal a more nuanced picture of its effectiveness. We demonstrate two things: - The effectiveness of Chain-of-Thought prompting can vary greatly depending on the type of task and model. For non-reasoning models, CoT generally improves average performance by a small amount, particularly if the model does not inherently engage in step-by-step processing by default. However, CoT can introduce more variability in answers, sometimes triggering occasional errors in questions the model would otherwise get right. We also found that many recent models perform some form of CoT reasoning even if not asked; for these models, a request to perform CoT had little impact. Performing CoT generally requires far more tokens (increasing cost and time) than direct answers. - For models designed with explicit reasoning capabilities, CoT prompting often results in only marginal, if any, gains in answer accuracy. However, it significantly increases the time and tokens needed to generate a response.
Problem

Research questions and friction points this paper is trying to address.

Investigates varying effectiveness of Chain-of-Thought prompting across tasks and models
Examines increased answer variability and occasional errors from Chain-of-Thought prompting
Assesses marginal accuracy gains versus higher costs in reasoning-enabled models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Investigates Chain-of-Thought prompting effectiveness
Shows CoT performance varies by task and model
Reveals CoT increases tokens and time costs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Lennart Meincke
WHU–Otto Beisheim School of Management
E
Ethan R. Mollick
Generative AI Labs, The Wharton School of Business, University of Pennsylvania
L
Lilach Mollick
Generative AI Labs, The Wharton School of Business, University of Pennsylvania
D
Dan Shapiro
Generative AI Labs, The Wharton School of Business, University of Pennsylvania, Glowforge