Extending Minimal Pairs with Ordinal Surprisal Curves and Entropy Across Applied Domains

📅 2026-03-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of the traditional minimal pair paradigm, which is constrained by binary grammaticality judgments and reliance on text generation, making it ill-suited for characterizing model uncertainty. The authors propose a generation-free evaluation framework that extends surprisal-based (negative log-probability) analysis from binary contrasts to multi-domain ordinal classification tasks. By constructing full surprisal curves, the method simultaneously captures both model preferences and uncertainty, and—novelty introduced here—incorporates entropy to distinguish between ambiguous and unambiguous samples. Evaluated across four application domains, the approach consistently yields interpretable signals: surprisal curves exhibit pronounced minima at the expected rating levels, thereby validating the method’s effectiveness and generalizability.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Relational Probabilistic ModelsNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successWeb Mining and Content Analysis: Models for Web evolution
📝 Abstract
The minimal pairs paradigm of comparing model probabilities for contrasting completions has proven useful for evaluating linguistic knowledge in language models, yet its application has largely been confined to binary grammaticality judgments over syntactic phenomena. Additionally, standard prompting-based evaluation requires expensive text generation, may elicit post-hoc rationalizations rather than model judgments, and discards information about model uncertainty. We address both limitations by extending surprisal-based evaluation from binary grammaticality contrasts to ordinal-scaled classification and scoring tasks across multiple domains. Rather than asking models to generate answers, we measure the information-theoretic "surprise" (negative log probability) they assign to each position on rating scales (e.g., 1-5 or 1-9), yielding full surprisal curves that reveal both the model's preferred response and its uncertainty via entropy. We explore this framework across four domains: social-ecological-technological systems classification, causal statement identification (binary and scaled), figurative language detection, and deductive qualitative coding. Across these domains, surprisal curves produce interpretable classification signals with clear minima near expected ordinal scale positions, and entropy over the completion tended to distinguish genuinely ambiguous items from easier items.
Problem

Research questions and friction points this paper is trying to address.

minimal pairs
ordinal classification
surprisal
model uncertainty
language model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

ordinal surprisal curves
entropy
minimal pairs
language model evaluation
uncertainty quantification
💼 Related Jobs
No related jobs found.
A
Andrew Katz
Department of Engineering Education, Virginia Tech