🤖 AI Summary
This study addresses the challenge of disentangling whether opinion shifts in large language models (LLMs) stem from relevant evidence or user pressure. To this end, it proposes a black-box evaluation framework that quantifies response variations across multiple scales, including probability, judgment, and action. By innovatively constructing conditional belief response profiles, the approach isolates these two influences, establishes the boundaries of diagnostic validity, and reveals the context-dependence of such evaluations. The framework is systematically validated through controlled variable analysis, synthetic data verification, and targeted ablation experiments. Experiments on models such as Qwen and Llama demonstrate that while surface-level response discrepancies are primarily driven by decoding strategies, certain internal belief patterns remain stable. These findings offer a novel paradigm for evaluating the robustness of LLMs.
📝 Abstract
A language model may revise the same proposition after receiving genuinely relevant evidence or after receiving directional user pressure that adds no relevant fact. The observable response shift alone therefore does not reveal which source drove the change. We introduce BeliefScope, a controlled black-box framework for separating these two sources of influence around a fixed target proposition. BeliefScope crosses Evidence and Pressure with factor-specific local controls and measures response changes through probability reports, categorical judgments, and action recommendations on channel-appropriate scales. To determine when these observable contrasts support reliable attribution, we evaluate the observation design under controlled synthetic conditions. Known-truth recovery and targeted ablations establish where Evidence- and Pressure-related effects can be separated, while semi-synthetic stress tests map how that recoverability changes as the observation process becomes noisier and more heterogeneous. Across a 36-family Qwen/Llama study, with targeted 12-family checks that also include Gemma3-12B, the resulting profiles show substantial evaluation-context dependence: broad model-level differences can change under matched controls, decoding, or response interfaces, while some narrower within-model patterns remain stable. Instruction interventions further show that reduced target-aligned Pressure following can reflect either stable resistance or movement in the opposite direction. BeliefScope summarizes these measurements as a conditional belief-response profile that keeps diagnostic effects tied to the evaluation conditions under which they are observed, together with explicit validity boundaries for that diagnosis.