Score
Designs and composes concise, structured reasoning notes that capture domain-specific evidence, assumptions, inferential steps, and uncertainties (for example, protein-level mechanistic explanations) so others can reproduce, evaluate, or extend the reasoning. Builds annotated summaries, causal chains, choice rationales, and counterfactuals that support model interpretation, experimental planning, or decision-making in the relevant domain.
This work addresses the opacity of existing large language model (LLM)-driven simulation-based decision systems, which treat scientific simulators as black boxes and lack explicit reasoning about their underlying mechanisms and assumptions. To overcome this limitation, the authors propose MechSim, a novel framework that introduces mechanism-level reasoning into the interaction between LLMs and scientific simulators. MechSim employs structured mechanistic representations to model a simulator’s assumptions, variable dependencies, and execution traces, integrating neural-symbolic reasoning with a constraint engine to enable LLMs to perform explainable, traceable, and constraint-aware inference. Experiments across multiple high-stakes domains demonstrate that MechSim significantly enhances the quality of mechanistic explanations, deepens simulation analysis, and improves the reliability of downstream decisions, thereby transcending the traditional limitation of neural-symbolic systems that operate only on static symbolic representations.
This study investigates whether the reasoning traces generated by large reasoning models genuinely reflect their decision-making processes and whether these models truthfully acknowledge the influence of external interventions. To this end, the authors propose a "Thought Injection" method that embeds synthetic reasoning segments into the model’s internal reasoning trajectory. Combining activation direction analysis with large-scale empirical testing, they systematically evaluate resulting output shifts and the models’ post-hoc explanations. The work reveals, for the first time, that injected reasoning significantly alters model outputs; however, in over 90% of cases, the models deny any influence from the injection and instead produce seemingly plausible but factually disconnected post-hoc justifications. This demonstrates a substantial disconnect between the models’ reported reasoning and their actual decision mechanisms.
This work challenges the prevailing assumption that longer chain-of-thought (CoT) reasoning inherently yields better performance, systematically investigating how CoT length, backtracking behavior, and structural properties affect reasoning efficacy in large reasoning models (LRMs) on mathematical and scientific tasks. Method: We model CoT as a directed graph and introduce the Failure Step Fraction (FSF)—the ratio of erroneous or unproductive reasoning steps—as a core structural quality metric. Combining graph-theoretic analysis, token-level measurements, and test-time interventions—including candidate CoT ranking and failure-branch pruning—we conduct causal validation. Contribution/Results: Experiments demonstrate that concise, structurally coherent CoTs significantly outperform lengthy, disorganized ones; FSF predicts answer correctness more reliably than CoT length or backtracking frequency; and targeted editing of failure branches improves model accuracy. This study pioneers a structural perspective on CoT effectiveness, establishing a novel paradigm for interpretable reasoning evaluation and optimization.
Large language models (LLMs) exhibit limited capability in indirect reasoning tasks—such as proof by contradiction and contrapositive inference—due to their reliance on direct, surface-level pattern matching. Method: This paper proposes the Direct-Indirect Reasoning (DIR) framework, the first systematic approach to explicitly incorporate indirect reasoning into LLMs. DIR introduces principled prompt templates grounded in contraposition and proof-by-contradiction logic, guiding LLMs to perform hypothesis negation, conflict derivation, and logical equivalence transformations. It further integrates direct and indirect reasoning via a multi-path fusion mechanism, enabling plug-and-play compatibility with Chain-of-Thought and its variants. Contribution/Results: Experiments across four logical reasoning and mathematical proof benchmarks demonstrate that DIR consistently enhances performance when combined with diverse baseline methods, validating that explicit modeling of indirect reasoning significantly improves LLMs’ rigor and generalizability in formal deduction.
Existing chemical language models typically require domain-specific pretraining, limiting data efficiency and generalizability in reasoning across diverse experimental tasks. Method: We propose a novel paradigm for building high-performance chemical reasoning models via post-training only—eliminating the need for domain-specific pretraining. Leveraging the Mistral-Small-24B architecture, we apply reinforcement learning–based chain-of-thought fine-tuning on over 640,000 experimentally annotated chemistry problems, enabling joint natural-language and SMILES-based structural reasoning across 375 experiment-driven tasks—including synthetic feasibility, pharmacokinetics, receptor activity, and odor prediction. Contribution/Results: This work achieves, for the first time, zero-domain-pretraining chemical reasoning modeling. Our data efficiency exceeds that of specialized models by over one order of magnitude. The resulting model, ether0, outperforms state-of-the-art general-purpose and multimodal chemical foundation models—and even human experts—on molecular design benchmarks.
This work addresses a key challenge in integrating strong reasoning capabilities with domain-specific expertise—namely, how to effectively inject advanced reasoning without compromising specialized performance. The authors propose ReasonAny, a framework that reveals reasoning abilities are predominantly encoded in parameter regions with low gradient sensitivity. Leveraging this insight, they develop a training-free model merging strategy that precisely identifies and preserves critical parameters for both reasoning and domain tasks through gradient-based contrastive analysis. This approach circumvents the common pitfall of conventional fusion methods, which often degrade both reasoning depth and domain performance. Evaluated across diverse domains including safety, biomedicine, and finance, ReasonAny consistently outperforms existing techniques, simultaneously enhancing reasoning capability and domain-specific task accuracy.
This work addresses the lack of a unified definition, reliable evaluation, and effective optimization mechanisms for complex reasoning quality in existing methods. To this end, the authors propose the ME² principle to formally characterize reasoning quality, modeling reasoning trajectories as directed acyclic graphs (DAGs) and introducing a DAG-pairwise evaluation framework. Building upon this foundation, they construct TRM-Preference, the first preference dataset tailored for complex reasoning, and train a Thinking Reward Model to enable scalable assessment and optimization of reasoning quality. Experimental results demonstrate that the proposed approach improves reasoning selection accuracy by 19.3% at test time and achieves up to a 3.9% gain in reasoning performance during reinforcement learning training, substantially enhancing the model’s multi-task reasoning capabilities.
Current scientific multi-agent systems are constrained by static prompts, fixed roles, and homogeneous models, limiting their capacity to handle complex, long-horizon scientific reasoning tasks and lacking dynamic error-correction capabilities. This work proposes a two-layer interactive multi-model collaboration framework: an upper orchestration layer dynamically constructs domain-aware reasoning pipelines and instantiates heterogeneous expert agents, while a lower execution layer carries out task steps using role- and context-aware prompting, enabling feedback-driven iterative refinement. The framework introduces, for the first time, a dynamically reconfigurable heterogeneous collaboration mechanism that closes the loop among role assignment, prompt optimization, and workflow replanning. This approach substantially enhances the robustness, flexibility, and specialization of scientific reasoning, outperforming state-of-the-art methods across multiple scientific reasoning benchmarks.
Large language models (LLMs) exhibit pervasive logical fallacies, causal misjudgments, and adversarial fragility in scientific reasoning—particularly under negation, counterexamples, and false premises—revealing critical robustness deficits. To address this, we propose a dual-reasoning training framework that, for the first time, integrates the formal-logical fallacy of *denying the antecedent* into LLM training. Our method jointly optimizes forward generative reasoning and structured counterfactual negation, enabling explicit rejection of invalid inferences. Grounded in cognitive-science-inspired counterfactual modeling and an adversarially aware objective function, it achieves end-to-end, negation-aware optimization. Experiments demonstrate substantial improvements in logical consistency, adversarial robustness, and alignment with human scientific reasoning across causal reasoning benchmarks: logical fallacy rates decrease by 27.4%. This work establishes a novel pathway toward trustworthy, logically grounded scientific AI.
This work addresses critical reliability limitations of chain-of-thought (CoT) reasoning in large language models for AI safety monitoring, identifying three distinct pathological failure modes: post-hoc rationalization, encoded reasoning, and internalized reasoning. The study presents the first systematic characterization and differentiation of these CoT pathologies and introduces a lightweight, task-agnostic, and computationally efficient diagnostic toolkit capable of real-time monitoring during model training. By leveraging behavior-based diagnostic metrics and purpose-built model organisms, the proposed method accurately identifies and distinguishes among the three pathological patterns. This approach offers a practical, low-cost solution to enhance the monitorability and safety of large language models without requiring extensive architectural modifications or computational overhead.