🤖 AI Summary
This study investigates whether large language models (LLMs) exhibit systematic biases toward either critical acclaim or popular appeal when evaluating films. Using a benchmark dataset of 200 movies, the authors conduct pairwise forced-choice experiments across eight LLMs from four model families, analyzing cultural preferences through Bradley–Terry preference modeling and nested ordinary least squares (OLS) regressions while controlling for public visibility and mainstream acceptance. The findings reveal, for the first time, a consistent “critic-oriented” bias across all models, which intensifies with increasing model scale. All LLMs significantly favor films with high critical recognition but low commercial success, whereas the advantage of films possessing dual legitimacy—both critical and popular—vanishes once visibility is controlled. Furthermore, the framing of prompts substantially influences recommendations, indicating that LLMs’ cultural preferences are highly sensitive to input formulation.
📝 Abstract
Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over commercially successful yet critically unrecognized ones. This pattern grows with model scale within each family. In addition, nested OLS regression analyses show that evaluative orientation, public visibility, and popular reception distinctly help explain preferences. Adjusting for public visibility reverses the models' preference for dual-legitimacy films over critical acclaim-only films, while additionally accounting for popular reception attenuates much of the disadvantage of films with commercial success only. Finally, evaluative and recommendation-oriented prompt framings produce divergent rankings, suggesting that critical acclaim orientation may manifest indirectly in real-world LLM deployments.