🤖 AI Summary
This study addresses the challenge of reliably capturing narrative themes and their dynamic evolution in small-scale poetic corpora, where traditional topic models often yield unstable results. To overcome this limitation, the authors propose a lightweight hybrid framework that integrates unsupervised Latent Dirichlet Allocation (LDA) with supervised sparse Partial Least Squares Discriminant Analysis (sPLS-DA). By incorporating multi-seed consensus strategies and narrative hub analysis—and deliberately filtering out prosodic and other surface-level linguistic features—the approach enables a computationally rigorous close reading of *Eugene Onegin*. The method substantially enhances topic stability and literary interpretability within limited corpora, successfully identifying five coherent themes that align meaningfully with the poem’s emotional trajectory and narrative arc. This work thus establishes a transparent, reproducible paradigm for computational analysis of densely layered literary texts.
📝 Abstract
This study presents a hybrid topic modelling framework for computational literary analysis that integrates Latent Dirichlet Allocation (LDA) with sparse Partial Least Squares Discriminant Analysis (sPLS-DA) to model thematic structure and longitudinal dynamics in narrative poetry. As a case study, we analyse Evgenij Onegin-Aleksandr S. Pushkin's novel in verse-using an Italian translation, testing whether unsupervised and supervised lexical structures converge in a small-corpus setting. The poetic text is segmented into thirty-five documents of lemmatised content words, from which five stable and interpretable topics emerge. To address small-corpus instability, a multi-seed consensus protocol is adopted. Using sPLS-DA as a supervised probe enhances interpretability by identifying lexical markers that refine each theme. Narrative hubs-groups of contiguous stanzas marking key episodes-extend the bag-of-words approach to the narrative level, revealing how thematic mixtures align with the poem's emotional and structural arc. Rather than replacing traditional literary interpretation, the proposed framework offers a computational form of close reading, illustrating how lightweight probabilistic models can yield reproducible thematic maps of complex poetic narratives, even when stylistic features such as metre, phonology, or native morphology are abstracted away. Despite relying on a single lemmatised translation, the approach provides a transparent methodological template applicable to other high-density literary texts in comparative studies.