Language models can notice an impossible engineering problem yet still report it as solved

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the phenomenon whereby large language models (LLMs) can identify physical contradictions in engineering problems yet frequently misreport unsolvable instances as resolved. To investigate this “knowing but not correcting” behavior, we evaluate 14 LLMs on 30 mechanics problems using a dual-verification mechanism, comparative experiments, and statistical significance analysis, with multidimensional assessment combining AI-based scoring and numerical validation. Results indicate that prompt optimization improves defect rejection rates but degrades problem-solving capability, revealing an inconsistency between model cognition and reporting. Accordingly, this work proposes a novel evaluation framework that independently assesses defect identification and final-state reporting, offering a new paradigm for evaluating LLM reliability.
📝 Abstract
Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of their flawed counterparts. Each reply required a"solved"or"cannot solve"status; rejection meant"cannot solve"or withholding an answer. The initial prompts did not warn that problems could be flawed. Across three recent models, 12 of 90 replies failed to reject a flawed problem. In 11 of these replies, the model stated the flaw, answered a corrected problem and still reported the original as"solved", according to artificial intelligence raters and numerical checks. We later retested four models from one provider, offering"flawed"instead of"cannot solve"and asking them to name and explain the defect. Three models showed statistically significant increases in rejection, but valid-problem solving fell in three. Evaluations therefore need to score both versions and distinguish flaw recognition from the reported status.
Problem

Research questions and friction points this paper is trying to address.

language models
engineering calculations
impossible problems
flaw recognition
evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language Models
Engineering Problem Solving
Flaw Recognition
Evaluation Methodology
Prompt Engineering
🔎 Similar Papers