🤖 AI Summary
This study addresses the challenge of retrieving biblical intertextuality in literary research, where allusions are often obscured by rewriting, indirect references, and translation. To this end, we construct a benchmark dataset and evaluate multiple retrieval models on the works of Karen Blixen. Methodologically, we propose conceptualizing retrieval models as “expert co-readers” to transcend the limitations of traditional annotation. Our approach employs TF-IDF, BM25, and multilingual as well as Danish sentence encoders, optimized through hard negative mining fine-tuning. Experimental results demonstrate that fine-tuning improves Recall@10 to 0.508 and doubles allusion retrieval performance. These findings effectively validate the potential of human-machine collaboration for uncovering deeper scholarly value in literary analysis.
📝 Abstract
Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Tales. Drawing on the commentary to a critical edition, we construct a benchmark of 189 annotated references and evaluate retrieval against all 31,170 verses of historically plausible Danish Old and New Testament translations. We compare TF-IDF and BM25 with multilingual and Danish sentence encoders, examine the effect of linguistic normalization, and fine-tune a Danish encoder using hard negatives and five-fold cross-validation. We analyze performance across automatically derived lexical-overlap strata representing quotations, paraphrases, and allusions. Linguistically normalized BM25 provides a strong zero-shot baseline, attaining an overall R@10 of 0.365 and retrieving every quotation within its ten highest-ranked verses. The best zero-shot dense model achieves a comparable overall score of 0.360 while performing better on allusions. Fine-tuning DFM-large raises its overall R@10 from 0.265 to 0.508 and more than doubles its performance on allusions, from 0.138 to 0.339. However, evaluation against editorial annotations alone understates the model's scholarly usefulness: a literary scholar judged seven of 30 selected rank-one predictions counted as false positives to be meaningful additional references. These findings show both the potential and the epistemic limits of computational intertextual retrieval. Rather than treating scholarly annotations as exhaustive or model outputs as discoveries, we propose retrieval models as heuristic co-readers that recover documented references and generate candidates for expert-led close reading.