🤖 AI Summary
This study addresses the challenge of effectively evaluating conformity behavior in large language models (LLMs) on open-ended tasks, where response quality is inherently subjective and lacks definitive ground-truth labels. The authors propose a systematic experimental protocol that disentangles re-answering effects, content exposure, peer presentation residuals, and evaluator context sensitivity. Their analysis reveals, for the first time, that evaluators’ own stances significantly influence scoring and that erroneous peer inputs degrade response quality. Using four open-source LLMs and three benchmark datasets, the study combines blind and informed evaluations, with GPT-4o and GPT-5.4-mini serving as auditing evaluators. Results demonstrate that fully incorrect peer inputs yield the lowest generation quality, evaluators exhibit heterogeneous responses to peer information, and uncalibrated anchors induce rating-scale instability—highlighting the critical importance of anchor calibration for reliable evaluation.
📝 Abstract
Prior work on LLM conformity largely measures discrete answer flips under verifiable labels. Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly. We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled peer-presentation residual, and directional judge sensitivity to visible peer context. Across four open-weight generators and three benchmarks, all-wrong peer input produces the lowest-quality revisions in every generator-dataset cell. Blind and informed ratings of identical answers also differ by evaluator: one judge shifts toward the peer-endorsed position, two shift away, one is approximately neutral, and GPT-4o and GPT-5.4-mini audits are likewise non-neutral. Finally, an anchor audit shows that terse correct anchors can be misread often enough to destabilize the latent scale unless calibration is checked explicitly. These results support four conclusions: flip rates are insufficient as a complete measure of open-ended conformity, wrong peers harm open-ended revision, evaluators are not neutral, and anchor calibration is necessary.