🤖 AI Summary
Large language models (LLMs) exhibit insufficient semantic alignment with human reasoning paths in chain-of-thought (CoT) inference, and their consistency degrades significantly over longer reasoning chains. Method: This paper proposes a semantic alignment evaluation framework that quantifies alignment via a novel alignment score; identifies two-hop reasoning chains as optimal for alignment through error-type analysis; and discovers four categories of reasoning biases that degrade performance. Building on these insights, we design Semantic Consistency Optimization Sampling (SCOS), a dynamic sampling method that selects intermediate reasoning steps with high semantic alignment. Contribution/Results: Experiments show that SCOS improves the average alignment score by 29.84% on three-hop reasoning tasks, substantially enhancing the semantic consistency between LLM-generated CoT traces and human logical reasoning paths. This work establishes a new paradigm for improving the interpretability and reliability of CoT reasoning in LLMs.
📝 Abstract
This paper presents a framework for evaluating and optimizing reasoning consistency in Large Language Models (LLMs) via a new metric, the Alignment Score, which quantifies the semantic alignment between model-generated reasoning chains and human-written reference chains in Chain-of-Thought (CoT) reasoning. Empirically, we find that 2-hop reasoning chains achieve the highest Alignment Score. To explain this phenomenon, we define four key error types: logical disconnection, thematic shift, redundant reasoning, and causal reversal, and show how each contributes to the degradation of the Alignment Score. Building on this analysis, we further propose Semantic Consistency Optimization Sampling (SCOS), a method that samples and favors chains with minimal alignment errors, significantly improving Alignment Scores by an average of 29.84% with longer reasoning chains, such as in 3-hop tasks.