Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

📅 2026-08-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the susceptibility of large language models (LLMs) to superficial stylistic cues when evaluating scientific creativity, which often undermines their ability to discern genuine scientific merit. To tackle this issue, the authors introduce SciStyleBench—the first framework specifically designed for diagnosing and mitigating style bias in this context. It comprises a three-stage controlled evaluation environment, a novel quantitative metric termed the Style Bias Index, and a plug-and-play style–content disentanglement module, SciStyleExtractor. Experimental results demonstrate that the proposed approach reduces the Style Bias Index from 0.566 to 0.501, while improving substantive recognition accuracy and adversarial win rate to 0.759 and 0.899, respectively, thereby significantly enhancing the robustness and fairness of LLM-based scientific assessment.
📝 Abstract
However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-Judge
stylistic bias
scientific substance
idea evaluation
style vs. substance
Innovation

Methods, ideas, or system contributions that make the work stand out.

stylistic bias
LLM-as-Judge
scientific idea evaluation
style-content disentanglement
evaluation robustness