🤖 AI Summary
This study addresses the lack of answer-availability awareness in large language models (LLMs) for Arctic science question answering, which frequently leads to unsupported responses. We construct the ArcticQA dataset with a paired benchmark and propose a joint evaluation framework integrating abstention frequency and response sensitivity. Through automated evidence verification, paired experimental design, and high-reasoning-effort testing, our approach effectively distinguishes baseline bias from genuine refusal capability. Experiments reveal significant differences in abstention rates across models; notably, replacing correct answers with incorrect ones increases the average abstention rate by 5.05%, confirming the models' latent sensitivity to answer availability. This work establishes a new paradigm for evaluating trustworthy refusal behavior in LLMs.
📝 Abstract
Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.