🤖 AI Summary
This study addresses the logical inconsistencies arising when probabilistic outputs of classification decision models violate axioms, noting that existing calibration methods fail to guarantee such consistency. To tackle this, the authors propose a label-free logical consistency testing framework that evaluates model probabilistic coherence through logically interrelated question answering. Utilizing the ChaosNLI and PubMedQA datasets, the work compares first-token and verbalized probabilities while conducting statistical significance analyses. The findings reveal substantial probability violations—such as mutually exclusive label probabilities summing beyond one—in models including Jev and Qwen, distinguishing their specific violation patterns and underlying causes. Furthermore, this research demonstrates that conventional evaluation metrics cannot capture these fundamental logical flaws, thereby establishing a novel paradigm for assessing model reliability.
📝 Abstract
Typed-decision models such as TypeSafe's Jev answer a declared yes/no or multiple-choice question about a state with a probability instead of text, and their evaluations report accuracy and calibration. Neither requires that the probabilities a model gives to logically related questions fit together item by item, which is what a system that acts on those probabilities needs. We test this property, coherence, with a battery of logically linked questions that needs no labels. For 160 items from ChaosNLI and PubMedQA, each with three mutually exclusive labels, we ask whether the label is X, whether it is not X, whether it is one of the other two labels, and which label applies. On 480 negation pairs, Jev's probabilities for"the label is X"and"the label is not X"miss summing to one by 0.064 on average (95% CI 0.055 to 0.072). Qwen3.8-27B, run from its official BF16 weights, misses by 0.293 with first-token probabilities and by 0.122 with verbalized probabilities. The gap to the first-token readout persists on pairs where both systems give similar probabilities, without double-negation labels, and after averaging Jev's repeated calls. Jev is not coherent either: its violations are about five times its repeat noise, and it over-endorses statements about single labels, so that its three single-label probabilities sum to 1.14 on average. The two systems also fail differently. Qwen3.8-27B's first-token readout under-endorses the complement of a label whether or not the question contains"not", rejecting both a statement and its negation in 196 of 480 pairs, and it does not become more coherent where it is more confident, whereas Jev's violations concentrate where its answer is uncertain. Because the checks need no labels, they expose biases that appear only when question forms are compared, and inconsistencies within items.