Likelihood Ranking doesn't Scale Like Prompting in LLMs

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the prevalent misconception in large language model evaluation that likelihood ranking and prompted answering are interchangeable. Using a controlled experimental design across 95 models and 10 datasets, we construct declarative sentences to compute likelihoods and systematically compare the scaling behaviors of likelihood ranking versus task-conditioned prompted answering. Our analysis reveals that likelihood preferences remain stable under model scaling, whereas prompted answering improves significantly with scale, demonstrating that these two paradigms probe distinct dimensions of model capability and exhibit systematic divergence. Accordingly, this work establishes that likelihood ranking and prompted answering should be employed as complementary rather than substitutable evaluation methodologies, clarifying their relationship and offering principled guidance for robust model assessment.
πŸ“ Abstract
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Likelihood Ranking
Prompting
Multiple-Choice QA
Model Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Likelihood Ranking
Prompting
LLM Evaluation
Multiple-Choice QA
Scaling
πŸ’Ό Related Jobs
No related jobs found.
A
Alessandro Bondielli
CoLingLab, Department of Philology, Literature and Linguistics, University of Pisa; Department of Computer Science, University of Pisa
L
Lucia Passaro
CoLingLab, Department of Philology, Literature and Linguistics, University of Pisa; Department of Computer Science, University of Pisa
Davide Bacciu
Davide Bacciu
Professor @ UniversitΓ  di Pisa - Founder & Chief Scientist @ ContinualIST.com
Deep LearningGenerative ModelsGraph Representation LearningContinual LearningPervasive AI
A
Alessandro Lenci
CoLingLab, Department of Philology, Literature and Linguistics, University of Pisa