๐ค AI Summary
This study addresses the frequent misjudgments in existing model consistency evaluations caused by conflating ambiguity with indifference. We propose a resampling-based, high-specificity contradiction detection algorithm and construct a 175-item disambiguated question set, establishing a new criterion thatๅคๅฎ inconsistency only when contradictions are statistically significant. This approach effectively eliminates ambiguity interference to precisely quantify model inconsistency. Our findings reveal that narrowly fine-tuned models exhibit severe pathologies, such as identity confusion, which fundamentally constrain their validity for alignment behavior research.
๐ Abstract
A large body of research measures model coherence based on output variance without adequately considering competing causes. We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either. We then measure incoherence in terms of contradictions when resampling answers to the same question. In contrast to other methods our metric has high specificity, and only ranks models as incoherent when the issues are glaring. Even so, we find narrow finetunes score poorly. Inspecting inconsistencies flagged by our method, we find that model organisms from the literature display severe issues such as identity conflation, introspection failures and rationalizations. These findings suggest that the pathologies induced by narrow finetuning may limit what these models can tell us about coherent misaligned behaviour.