🤖 AI Summary
This study addresses the limitation of existing factuality evaluations for large language models (LLMs), which typically focus on isolated judgments while neglecting the persistence of beliefs within a knowledge system. To this end, this work proposes a hierarchical belief stability metric that leverages internal model representations to estimate conditional probabilities, shifting the evaluation from singular belief strength assessment to relational stability analysis that measures a belief's consistency with other cognitive commitments. Empirical evaluations across twelve LLMs and three domains demonstrate that in 83.3% of the experimental settings, beliefs exhibiting low stability are more susceptible to fluctuation under conversational challenges. This work effectively extends the dimensions of model reliability evaluation.
📝 Abstract
Large language models (LLMs) increasingly mediate how people access and reason with information, yet factual reliability is usually evaluated one judgment at a time. We introduce graded belief stability, a relational measure of how well a belief persists within an LLM's broader belief system. Unlike individual belief probability, it asks whether support for a claim persists when that claim is considered alongside the model's other epistemic commitments. We operationalize this idea with a Direct Conditional estimator that uses internal model representations to estimate conditional belief probabilities. Across 12 LLMs and three domains, lower-stability beliefs exhibit greater mean behavioral movement under conversational challenge in 83.3% of model-domain settings after matching on individual belief probability. Graded belief stability therefore extends reliability assessment beyond how strongly an LLM supports a claim to how robustly that belief is supported within its broader system of beliefs.