🤖 AI Summary
This study addresses the diminishing discriminative power of multiple-choice benchmarks due to performance saturation and the inefficiency of manual question construction. We propose AnswerPool, a method that eliminates the need for generating new questions by merging options from items sharing a common context into composite tasks, requiring models to perform answer assignment. This substantially increases evaluation difficulty while simultaneously assessing abstention capabilities. Experiments across eight benchmarks and eighteen models demonstrate that AnswerPool significantly elevates task difficulty for all evaluated models, effectively quantifies scoring biases introduced by elimination strategies, reveals weaker models' reliance on such heuristics, and exposes the inability of most open-source models to correctly identify unanswerable questions.
📝 Abstract
Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question's wrong options are written for that question alone, so a model can score by eliminating a few options. We propose AnswerPool: take $N$ questions that share a context, pool all their options into one list, and ask the model to assign every question its answer. No item is written and no label changes. The chance of guessing a group right falls from $10^{-3}$ to $5\times10^{-7}$ for five four-option questions, and a model that recognizes its answers keeps its multiple-choice score, so the accuracy lost to pooling measures the credit the format gave for elimination. Deleting answers from the pool makes questions unanswerable with exact ground truth, so abstention is scored in the same pass. Across eight text, image, and video benchmarks and eighteen models, pooling is harder for every model, the elimination credit is largest for the weakest models, and seven of eight open-weight models answer 87 to 100% of unanswerable questions.