🤖 AI Summary
This study addresses the phenomenon wherein language models frequently generate correct candidate answers yet struggle to ultimately select them. To investigate this, the authors decouple factual recall into three distinct stages—generation, ranking, and selection—revealing a dissociation between what models “know” and what they can “select.” By adopting an explicit verification strategy based on P(True) and comparing it against an average log-likelihood baseline across models such as Gemma and Qwen3, this work demonstrates that explicit verification outperforms implicit generation preferences. It further highlights the substantial influence of evaluation metrics on reported outcomes. Experimental results indicate that the proposed approach improves ranking AUROC by 0.08–0.12 and increases majority accuracy by approximately 5%, maintaining advantages even over strong baselines, although performance remains constrained by entity accessibility.
📝 Abstract
Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across three model families, but say little about whether an available correct answer will ultimately be selected. Explicit verification with $P(\mathrm{True})$ improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, with AUROC gains of $0.08$--$0.12$. In a prospectively defined Gemma cohort, verification raises plurality accuracy by about $5$ points, and still gains about $2$ points over chat-template likelihood, a stronger generative baseline. The advantage is strongest for relations with common-answer priors and depends on access to the entity; masking the entity removes the ranking advantage in larger Qwen models. Finally, the measured benefit depends on how correctness is defined: recall-oriented reference matching can credit option lists favored by likelihood and substantially understate the improvement seen under human semantic judgments.