Feedback Indices to Evaluate LLM Responses to Rebuttals for Multiple Choice Type Questions

📅 2026-01-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of systematic evaluation of large language models’ responses to user disagreement, particularly their susceptibility to undesirable conversational tendencies such as flattery or stubbornness. The authors propose a feedback metric framework based on fictitious response-rebuttal (FR) pairings, enabling the first quantitative assessment of model flattery and stubbornness in multiple-choice settings without ground-truth answers. This approach facilitates cross-model and cross-topic comparisons of dialogic behavior. Empirical evaluation on physics-related questions reveals that newer OpenAI models and higher reasoning configurations significantly reduce flattery, thereby demonstrating the framework’s effectiveness and generalizability.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Large Multimodal Models (LMMs)Humans and AI: Learning Human Values and Preferences

Application Category

User Modeling, Personalization and Recommendation: Metrics for user behavior and evaluating successSearch and Retrieval-Augmented AI: Web evaluation methodologies and metricsSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
We present a systematic framework of indices designed to characterize Large Language Model (LLM) responses when challenged with rebuttals during a chat. Assessing how LLMs respond to user dissent is crucial for understanding their reliability and behavior patterns, yet the complexity of human-LLM interactions makes systematic evaluation challenging. Our approach employs a fictitious-response rebuttal method that quantifies LLM behavior when presented with multiple-choice questions followed by deliberate challenges to their fictitious previous response. The indices are specifically designed to detect and measure what could be characterized as sycophantic behavior (excessive agreement with user challenges) or stubborn responses (rigid adherence to the fictitious response in the chat history) from LLMs. These metrics allow investigation of the relationships between sycophancy, stubbornness, and the model's actual mastery of the subject matter. We demonstrate the utility of these indices using two physics problems as test scenarios with various OpenAI models. The framework is intentionally generalizable to any multiple-choice format question, including on topics without universally accepted correct answers. Our results reveal measurable differences across OpenAI model generations, with trends indicating that newer models and those employing greater"Reasoning Effort"exhibit reduced sycophantic behavior. The FR pairing method combined with our proposed indices provides a practical, adaptable toolkit for systematically comparing LLM dialogue behaviors across different models and contexts.
Problem

Research questions and friction points this paper is trying to address.

LLM evaluation
sycophancy
stubbornness
rebuttal response
multiple-choice questions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fictitious-Response Rebuttal
Sycophancy Detection
Stubbornness Metrics
LLM Behavior Evaluation
Reasoning Effort
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
J. Dunlap
Portland State University, Portland, Oregon, United States
A
Anne-Simone Parent
University of Liège, Liège, Belgium
R
R. Widenhorn
Portland State University, Portland, Oregon, United States