Chasing Consistency: Quantifying and Optimizing Human-Model Alignment in Chain-of-Thought Reasoning

📅 2025-11-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Large language models (LLMs) exhibit insufficient semantic alignment with human reasoning paths in chain-of-thought (CoT) inference, and their consistency degrades significantly over longer reasoning chains. Method: This paper proposes a semantic alignment evaluation framework that quantifies alignment via a novel alignment score; identifies two-hop reasoning chains as optimal for alignment through error-type analysis; and discovers four categories of reasoning biases that degrade performance. Building on these insights, we design Semantic Consistency Optimization Sampling (SCOS), a dynamic sampling method that selects intermediate reasoning steps with high semantic alignment. Contribution/Results: Experiments show that SCOS improves the average alignment score by 29.84% on three-hop reasoning tasks, substantially enhancing the semantic consistency between LLM-generated CoT traces and human logical reasoning paths. This work establishes a new paradigm for improving the interpretability and reliability of CoT reasoning in LLMs.

Technology Category

Cognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningSearch and Optimization: Learning to SearchMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved information
📝 Abstract
This paper presents a framework for evaluating and optimizing reasoning consistency in Large Language Models (LLMs) via a new metric, the Alignment Score, which quantifies the semantic alignment between model-generated reasoning chains and human-written reference chains in Chain-of-Thought (CoT) reasoning. Empirically, we find that 2-hop reasoning chains achieve the highest Alignment Score. To explain this phenomenon, we define four key error types: logical disconnection, thematic shift, redundant reasoning, and causal reversal, and show how each contributes to the degradation of the Alignment Score. Building on this analysis, we further propose Semantic Consistency Optimization Sampling (SCOS), a method that samples and favors chains with minimal alignment errors, significantly improving Alignment Scores by an average of 29.84% with longer reasoning chains, such as in 3-hop tasks.
Problem

Research questions and friction points this paper is trying to address.

Quantifying semantic alignment between model and human reasoning chains
Identifying key error types that degrade reasoning consistency in LLMs
Optimizing sampling to minimize alignment errors in longer reasoning chains
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces Alignment Score metric for human-model reasoning alignment
Defines four error types degrading semantic consistency in reasoning
Proposes SCOS method to optimize chains by minimizing alignment errors
B
Boxuan Wang
School of Computer Science and Informatics, University of Liverpool
Z
Zhuoyun Li
School of Computer Science and Informatics, University of Liverpool
X
Xinmiao Huang
School of Computer Science and Informatics, University of Liverpool
Xiaowei Huang
Xiaowei Huang
Professor of Computer Science, University of Liverpool
AI Safety and SecurityVerificationTrustworthy AIFormal MethodsExplainable AI
Y
Yi Dong
School of Computer Science and Informatics, University of Liverpool