A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

๐Ÿ“… 2026-07-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Current evaluations of mathematical reasoning rely solely on answer matching, failing to detect invalid reasoning chains that happen to yield correct answersโ€”a phenomenon this work terms the โ€œreasoning-answer consistency gap.โ€ To address this, we propose RAFS, a novel scoring mechanism that formally quantifies this gap through reference-free, instance-level diagnostic evaluation of reasoning trajectories. RAFS assesses local plausibility, support for the final answer, and stability under resampling and counterfactual interventions, integrating step validity, reasoning-to-answer entailment, counterfactual sensitivity, and conditional stability. Employing a pre-registered confirmatory study design, experiments on GSM8K and MATH demonstrate that RAFS provides auditable early warnings of silent failures and quantifies the trade-off between computational overhead and abstention, offering a complementary diagnostic tool for evaluating mathematical reasoning.
๐Ÿ“ Abstract
Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answer, while a valid calculation can be followed by a transcription error. We call this mismatch the reasoning answer consistency gap. This framework paper introduces the Reasoning Answer Faithfulness Score (RAFS), a reference free, instance level diagnostic of whether an emitted mathematical trace is locally credible, supports its answer, and is stable under resampling and targeted counterfactual interventions. RAFS combines step validity, reasoning to answer entailment and counterfactual sensitivity, answer consensus, and conditional reasoning stability. It evaluates transcript level agreement, not a models private computation and not factual correctness outside the tested mathematical setting. We retain a preregistered, results blind confirmatory study on GSM8K and MATH, with hypotheses, admissibility rules, calibration, and tests fixed before confirmatory outcomes are inspected. A separate feasibility pilot is specified to verify end to end execution and estimate interven tion coverage before that freeze numerical pilot claims are re ported only when trace level artifacts are available. We formalize four reasoning answer outcomes, justify the non compensatory aggregator, instantiate semantic trace distance, quantify compute and abstention tradeoffs, and define verifier independence and power analyses. RAFS is intended to complement mathematical answer accuracy with an auditable warning signal for silent reasoning failures and answer extraction errors
Problem

Research questions and friction points this paper is trying to address.

reasoning-answer consistency gap
silent reasoning failures
chain of thought evaluation
reference-free scoring
mathematical reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

reference-free evaluation
reasoning faithfulness
silent reasoning failures
counterfactual sensitivity
chain-of-thought diagnostics
๐Ÿ’ผ Related Jobs
No related jobs found.
V
Vivek Shukla
Allenhouse Institute of Technology
V
Varun Shukla
Allenhouse Institute of Technology
A
Atul
Allenhouse Institute of Technology
D
Divya Mishra
Allenhouse Institute of Technology
M
Mehul Kumar Das
Allenhouse Institute of Technology