🤖 AI Summary
This study addresses the absence of explicit measurement and control over perturbation strength in evaluating the self-consistency of large language model (LLM) explanations. We propose a unified framework for quantifying perturbation intensity across both inputs and chains-of-thought (CoT). The core innovation lies in introducing an LLM-as-a-judge-based perturbation metric, enabling unified scaling and fair comparison across heterogeneous perturbation types, alongside embedding- and probability-based baselines for systematic evaluation. Experimental results demonstrate that the proposed metric significantly outperforms conventional approaches. Crucially, our analysis reveals that input perturbations exert a substantially greater impact on explanation consistency than CoT perturbations. Furthermore, this work establishes the principle of intra-type comparison, offering a novel paradigm for assessing the reliability of LLM-generated explanations.
📝 Abstract
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.