🤖 AI Summary
This work proposes a stylometric approach inspired by genome-wide association studies (GWAS), treating words as “genetic variants” and authorship as the “phenotype.” By applying logistic regression combined with multiple testing correction, the method identifies author-specific lexical markers that exhibit statistical significance in textual data. It represents the first adaptation of the GWAS paradigm to stylometric analysis, offering both statistical rigor and interpretability of results. Experimental validation across multilingual corpora—including English, German, and Russian—demonstrates the method’s ability to reliably detect stable, author-unique lexical features, thereby confirming its effectiveness and cross-lingual generalizability.
📝 Abstract
This short paper introduces a stylometric interpretation method inspired by genome-wide association studies (GWAS). Each "gene" token's association with "phenotype" authorship is tested using logistic regression with multiple-comparison correction. Applied to English, German, and Russian corpora, the method detects statistically significant lexical markers distinctive of individual authors.