π€ AI Summary
This work addresses the limitation of existing cultural benchmarks, which predominantly assess factual knowledge while neglecting deeper reasoning capabilities such as explaining, substantiating, and revising cultural references. Using literary interpretation as the evaluation context, this study proposes an evidence-centered benchmark that systematically examines modelsβ deep cultural understanding through the cross-contextual identification and reconstruction of cultural references. Methodologically, it integrates literary data analysis, contextual resources, and expert feedback mechanisms. The framework is validated using Danish literature case studies to evaluate modelsβ cultural robustness and interpretive depth while preserving legitimate scholarly disagreement. Ultimately, this research establishes an evaluation paradigm that transcends conventional metrics, advancing AI development toward systems capable of sophisticated cultural reasoning.
π Abstract
How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.