TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

๐Ÿ“… 2026-05-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Current evaluations of large language model reasoning predominantly rely on final answer accuracy or superficial statistical features, which inadequately capture the quality of reasoning processes in open-ended outputs. This work proposes TRACE, a novel metric that, for the first time, integrates Toulminโ€™s argumentation model with Flavellโ€™s metacognitive framework to perform fine-grained structural analysis of chain-of-thought reasoning, thereby enabling quantitative assessment of the intrinsic quality of reasoning construction. TRACE can serve as a reward signal in reinforcement learning. Experiments across seven models and 26.3K question-answer pairs demonstrate that TRACE exhibits strong correlation with benchmark accuracy (r = 0.74) and significantly outperforms reinforcement learning baselines that rely solely on answer accuracy.
๐Ÿ“ Abstract
Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-answer accuracy or surface-level statistics, leaving the reasoning process itself unexamined. We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. Rather than judging outcomes, TRACE inspects how arguments are constructed by integrating Toulmin's argumentation theory with Flavell's metacognitive framework to assess reasoning structure. Experiments on 26.3K QA samples across 7 reasoning models show strong correlation with benchmark accuracy (r=0.74). Furthermore, TRACE is effective as a reinforcement learning reward signal, outperforming accuracy-only baselines. Together, these results indicate that logically sound reasoning leads to higher-quality answers. TRACE thus serves as a complementary metric for evaluating open-ended outputs. Code is available at https://github.com/hyyangkisti/trace.
Problem

Research questions and friction points this paper is trying to address.

reasoning evaluation
large language models
Chain-of-Thought
open-ended outputs
argumentation structure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Toulmin argumentation
Chain-of-Thought evaluation
reasoning assessment
metacognitive framework
LLM evaluation
Y
Yundong Kim
Applied Agent Research Center, Korea Institute of Science and Technology Information (KISTI), Republic of Korea; Department of Computer Science and Engineering, University of Seoul, Republic of Korea
H
Heyoung Yang
Applied Agent Research Center, Korea Institute of Science and Technology Information (KISTI), Republic of Korea