Score
Designs, builds, or analyzes mechanisms that measure and enforce semantic agreement between a model’s intermediate reasoning traces (thoughts, chains-of-thought) and its final answer, including scoring functions and learned reward models that quantify trace–answer alignment. Implements evaluation metrics and training or inference-time objectives (for example, consistency rewards or augmented RL objectives) to detect, score, and minimize thinking–answer inconsistencies during generation or policy updates.
Large language models (LLMs) have long relied on outcome-based reward modeling—evaluating only final answers—leading to insufficient interpretability and robustness in reasoning processes. Method: This paper systematically introduces the Process Reward Modeling (PRM) paradigm, shifting supervision from outcome-level to step- or trajectory-level reasoning evaluation. We establish a comprehensive methodology encompassing data construction, fine-grained reward modeling, test-time scaling, and RLHF integration. Contribution/Results: We empirically validate PRM across diverse domains—including mathematics, code generation, natural language, multimodal reasoning, and robotic agent tasks. To our knowledge, this is the first work to characterize the design space and core challenges of PRMs across multiple domains, releasing a cross-task benchmark, practical implementation guidelines, and open-source resources. Our framework provides foundational theoretical insights, actionable technical pathways, and empirical support for trustworthy reasoning alignment in LLMs.
The evaluation of step-by-step reasoning quality in large language models lacks standardized benchmarks, resulting in fragmented metrics and inconsistent evaluation practices. Method: We systematically survey over 60 reasoning evaluation methods and propose the first taxonomy for reasoning trace assessment, structured along four dimensions—factual consistency, validity, coherence, and practical utility. Through cross-dimension experiments, we empirically assess the transferability of evaluation models across these criteria and identify their feasibility boundaries. Contribution/Results: We introduce a standardized meta-evaluation benchmark and establish the first structured reasoning assessment framework, explicitly mapping evaluation metrics to underlying criteria. This framework clarifies conceptual relationships among existing methods, enables systematic comparison, and lays the foundation for rigorous, reproducible, and comparable reasoning evaluation. Our work bridges critical gaps between theoretical desiderata and practical assessment, advancing the standardization and scientific rigor of LLM reasoning evaluation.
This study investigates whether chain-of-thought (CoT) reasoning traces faithfully reflect a model’s actual internal decision-making process, thereby questioning their reliability as a supervisory and auditing mechanism. To this end, the authors propose a step-level Detect-Classify-Compare framework, integrating multidimensional validation techniques—including answer-commitment agents, Patchscopes, tuned-lens probes, causal ablation, truncation experiments, and donor contamination tests. Experiments across nine models and seven reasoning benchmarks reveal that, on average, only 61.9% of CoT steps align with the model’s internal computations; in 58% of misaligned cases, models generate redundant “reasoning” after the answer has already been determined—a phenomenon termed “hallucinated continuation.” Notably, stronger CoT performance correlates with lower temporal fidelity. This work provides the first systematic evidence of a fundamental disconnect between CoT traces and genuine reasoning dynamics, challenging the core assumption that CoT serves as a faithful reasoning log.
Large language models (LLMs) exhibit insufficient semantic alignment with human reasoning paths in chain-of-thought (CoT) inference, and their consistency degrades significantly over longer reasoning chains. Method: This paper proposes a semantic alignment evaluation framework that quantifies alignment via a novel alignment score; identifies two-hop reasoning chains as optimal for alignment through error-type analysis; and discovers four categories of reasoning biases that degrade performance. Building on these insights, we design Semantic Consistency Optimization Sampling (SCOS), a dynamic sampling method that selects intermediate reasoning steps with high semantic alignment. Contribution/Results: Experiments show that SCOS improves the average alignment score by 29.84% on three-hop reasoning tasks, substantially enhancing the semantic consistency between LLM-generated CoT traces and human logical reasoning paths. This work establishes a new paradigm for improving the interpretability and reliability of CoT reasoning in LLMs.
This work addresses the issue of “deceptive alignment” in existing reward models, where correct judgments are accompanied by flawed reasoning due to exclusive optimization for outcome accuracy, thereby limiting generalization in reinforcement learning from human feedback (RLHF). To mitigate this, the authors propose a “reasoning consistency” metric that quantifies the alignment between a model’s internal reasoning process and human judgment. This metric is combined with outcome accuracy to form a hybrid training signal for a generative reward model (GenRM). Notably, this approach introduces fine-grained reasoning consistency evaluation for the first time. Experimental results demonstrate that GenRM achieves state-of-the-art performance on RM-Bench (87.1%) and JudgeBench (82%), outperforming baselines by an average of 5%, and further improves creative writing performance by 7% on Arena Hard v2, significantly enhancing reasoning consistency.
Relying solely on the correctness of final answers as a reward often leads to unreliable reasoning processes and limited downstream utility. This work proposes TraceLift, a novel framework that introduces “executor-anchored rewards,” treating reasoning trajectories as intermediate artifacts intended for consumption by downstream executors. By jointly optimizing trajectory quality through a rule-based reasoning reward model and measurable improvements in executor performance, TraceLift ensures that generated traces are both logically sound and practically useful. To support fine-grained supervision, we construct the TRACELIFT-GROUPS dataset, comprising high-quality reference trajectories alongside their locally perturbed variants for the same problems. Experiments on mathematical and code generation tasks demonstrate that our approach significantly outperforms training strategies based exclusively on execution outcomes, confirming that effective reasoning trajectories must balance formal coherence with tangible benefits to downstream models.
This study investigates the impact of transferring chain-of-thought (CoT) reasoning from one large language model to another on the recipient model’s inference and generation mechanisms. By establishing a provider–recipient framework and employing techniques such as CoT prefix truncation, forced-answer versus free-generation comparisons, and multi-model, multi-benchmark evaluation, the work reveals that CoT transfer operates through multiple pathways—including answer extraction, reasoning scaffolding, and dependence on the recipient model’s inherent capabilities—rather than a single uniform mechanism. The authors propose using answer consistency in the absence of ground-truth labels as an early stopping signal for reasoning and demonstrate across benchmarks including AIME, MMLU-Pro, and ZebraLogic that partial CoT prompts can effectively guide subsequent reasoning and enhance performance.
Existing approaches struggle to detect distributed, non-local errors in large language model reasoning and rely heavily on the strong assumption of semantic faithfulness in chain-of-thought (CoT) outputs. This work proposes a diagnostic framework that dispenses with this assumption by analyzing dynamic structural changes in CoT reasoning processes, revealing for the first time that reasoning failures manifest as task-dependent structural anomalies. Through controlled experiments on Boolean satisfiability tasks, sentence function annotation, dynamic behavior analysis, and targeted prompt interventions, the method achieves a substantial improvement in error detection accuracy—increasing from 13.3% to 85% on Llama3-70B—and successfully corrects 84.6% of identified reasoning errors.
Current preference alignment methods for large language models predominantly rely on outcome-level rewards, which provide insufficient fine-grained guidance over the reasoning process and consequently lead to inadequate trajectory-level preference modeling. To address this limitation, this work proposes Thinking Checklist Reward (TCR), which introduces sample-specific thinking checklists to reformulate preference alignment as an evaluation of how well key reasoning considerations are covered throughout the inference trajectory. TCR further incorporates an exponential moving average (EMA) residual mechanism to effectively disentangle reasoning quality from final output quality, thereby isolating a “reasoning gain” that transcends outcome-based rewards. Experiments across five models spanning three model families demonstrate consistent improvements in alignment performance on multiple benchmarks, while ablation studies confirm the critical contributions of both checklist supervision and the EMA residual design.
This work addresses the prevalent yet often undetectable issue of logical inconsistency between reasoning and final answers in chain-of-thought (CoT) outputs generated by current AI systems during safety evaluations. The study is the first to formally distinguish between reasoning consistency and faithfulness, introducing a taxonomy encompassing six distinct types of inconsistency. To enable post-hoc detection without modifying model generation, the authors propose InspectScout—a reusable scanning method grounded in formal definitions, supported by a human-annotated benchmark, and implemented via an automated detection algorithm integrated into the inspect_evals framework. Experiments demonstrate that reasoning inconsistencies are widespread across four mainstream models and three safety-related tasks, and can be reliably identified; moreover, the patterns of such inconsistencies exhibit systematic variation across models.
Current evaluations of mathematical reasoning rely solely on answer matching, failing to detect invalid reasoning chains that happen to yield correct answers—a phenomenon this work terms the “reasoning-answer consistency gap.” To address this, we propose RAFS, a novel scoring mechanism that formally quantifies this gap through reference-free, instance-level diagnostic evaluation of reasoning trajectories. RAFS assesses local plausibility, support for the final answer, and stability under resampling and counterfactual interventions, integrating step validity, reasoning-to-answer entailment, counterfactual sensitivity, and conditional stability. Employing a pre-registered confirmatory study design, experiments on GSM8K and MATH demonstrate that RAFS provides auditable early warnings of silent failures and quantifies the trade-off between computational overhead and abstention, offering a complementary diagnostic tool for evaluating mathematical reasoning.
This study addresses a critical gap in current evaluation methods, which focus solely on final answer accuracy while overlooking how models identify and annotate biased content during reasoning, thereby creating blind spots in accountability assessment. To remedy this, the authors propose a fine-grained diagnostic framework based on reasoning traces that evaluates model behavior along two dimensions: bias sensitivity and bias acknowledgment—the latter being a novel metric introduced in this work to measure whether a model explicitly flags biased content within its Chain-of-Thought reasoning. By integrating human-defined surface indicator rules, the framework enables automated analysis of reasoning trajectories. Experiments on GSM8K reveal that while GPT-4o and Claude Sonnet 4 exhibit comparable bias sensitivity, their bias acknowledgment rates differ markedly at 13.0% and 75.0%, respectively, highlighting substantial disparities in responsible reasoning capabilities.