Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current scientific reasoning benchmarks focus solely on the correctness of final answers, failing to detect when large language models arrive at correct responses through invalid shortcuts such as enumeration or guessing. This work systematically uncovers and quantifies, for the first time, the prevalence of such “solution hacking” behavior among state-of-the-art models on challenging scientific tasks, revealing that up to 44.1% of ostensibly correct answers are actually obtained via these spurious strategies. To address this issue, we introduce an expert-informed anti-shortcut mechanism comprising an automatic discriminator and test-time instruction interventions, demonstrating its effectiveness across diverse difficulty levels and domains. Upon suppressing these shortcuts, model accuracy drops substantially, indicating that existing evaluations significantly overestimate models’ genuine reasoning capabilities.
📝 Abstract
Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necessarily demonstrate the reasoning capability targeted by the problem. We identify Solution Hacking, a failure mode in which an LLM reaches the correct answer through invalid shortcuts, such as numerical search, enumeration, guessing, or answer-first verification, without providing a valid task-targeted derivation. We systematically analyze this phenomenon across difficulty levels, scientific domains, and frontier models. Solution hacking increases sharply with benchmark difficulty, from 2.2\% on common problems to 28.3\% on Olympiad-level problems and 37.4\% on HLE. Moreover, 8.2\%-44.1\% of answers credited as correct across frontier models are identified as hacked solutions. We further develop expert-inspired anti-hacking strategies, including an automatic judge and a test-time instruction. The results show that suppressing shortcut behavior substantially reduces reported accuracy while having a smaller effect on correct and non-hacked accuracy. These findings reveal that answer-only evaluation can overestimate the scientific reasoning capabilities of frontier LLMs.
Problem

Research questions and friction points this paper is trying to address.

Solution Hacking
Scientific Reasoning
Large Language Models
Benchmark Evaluation
Reasoning Capability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Solution Hacking
scientific reasoning
large language models
answer-only evaluation
anti-hacking strategies
🔎 Similar Papers
No similar papers found.