🤖 AI Summary
Large language models (LLMs) exhibit pervasive logical fallacies, causal misjudgments, and adversarial fragility in scientific reasoning—particularly under negation, counterexamples, and false premises—revealing critical robustness deficits. To address this, we propose a dual-reasoning training framework that, for the first time, integrates the formal-logical fallacy of *denying the antecedent* into LLM training. Our method jointly optimizes forward generative reasoning and structured counterfactual negation, enabling explicit rejection of invalid inferences. Grounded in cognitive-science-inspired counterfactual modeling and an adversarially aware objective function, it achieves end-to-end, negation-aware optimization. Experiments demonstrate substantial improvements in logical consistency, adversarial robustness, and alignment with human scientific reasoning across causal reasoning benchmarks: logical fallacy rates decrease by 27.4%. This work establishes a novel pathway toward trustworthy, logically grounded scientific AI.
📝 Abstract
Large Language Models (LLMs) have transformed natural language processing and hold growing promise for advancing science, healthcare, and decision-making. Yet their training paradigms remain dominated by affirmation-based inference, akin to extit{modus ponens}, where accepted premises yield predicted consequents. While effective for generative fluency, this one-directional approach leaves models vulnerable to logical fallacies, adversarial manipulation, and failures in causal reasoning. This paper makes two contributions. First, it demonstrates how existing LLMs from major platforms exhibit systematic weaknesses when reasoning in scientific domains with negation, counterexamples, or faulty premises footnote{Code to recreate these experiments are at https://github.com/hannahdavidsoncollege-maker/ScientificReasoningForEnvironment-MedicineWithLLMs. Second, it introduces a dual-reasoning training framework that integrates affirmative generation with structured counterfactual denial. Grounded in formal logic, cognitive science, and adversarial training, this training paradigm formalizes a computational analogue of ``denying the antecedent'' as a mechanism for disconfirmation and robustness. By coupling generative synthesis with explicit negation-aware objectives, the framework enables models that not only affirm valid inferences but also reject invalid ones, yielding systems that are more resilient, interpretable, and aligned with human reasoning.