MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Mechanistic descriptions generated by large language models (LLMs) lack causal structural constraints, undermining their reliability. This work addresses this limitation by leveraging confluence-based qualitative physics to construct an auditable NLP benchmark and simulator. Through techniques including component state modeling, topological verification, and quantitative simulation-based behavioral validation, the proposed framework ensures via deterministic generation that reasoning processes strictly adhere to underlying structural and causal constraints. Experiments demonstrate that GPT-5.5 accuracy declines significantly from 76.1% to 38.0% as task complexity increases, revealing inherent limitations of existing models in ambiguity resolution and causal inference. These findings confirm the effectiveness of the proposed framework for reliably evaluating LLMs.
📝 Abstract
This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying structural and causal constraints. The benchmark tests whether answers preserve simulator-licensed ambiguity, quantified claims, episode-graph transition evidence, repairs, and trace-support judgments. Its 1,120 items are generated deterministically from admissible interpretation sets, component states, scenario restrictions, confluence constraints, and derivation steps across 18 catalog mechanisms and six task families. Each mechanism undergoes converter checks of structure and topology and behavioral checks against quantitative simulations. GPT-5.5 accuracy decreases as family-specific mechanistic complexity increases, from 76.1% in the lowest-complexity bucket (B1) to 38.0% in the highest-complexity bucket (B4). The negative association remains after controls for rendered-prompt and expected-answer length. These results show that qualitative simulators can support auditable NLP benchmarks for mechanistic inference.
Problem

Research questions and friction points this paper is trying to address.

mechanistic reasoning
qualitative physics
large language models
benchmark evaluation
causal constraints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Qualitative Physics
Mechanistic Reasoning
Confluence-based Simulation
Auditable Benchmark
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Danilo Gusicuma
Idiap Research Institute, Switzerland; École Polytechnique Fédérale de Lausanne (EPFL), Switzerland
André Freitas
André Freitas
University of Manchester | Idiap Research Institute
Artificial IntelligenceNatural Language ProcessingNatural Language InferenceNeuro-symbolic AI