🤖 AI Summary
This study addresses the challenge of capturing stylistic features in zero-shot authorship attribution due to the absence of supervised signals. It systematically compares large language model prompting strategies with embedding-based representations and proposes a two-stage framework, LISA. This method optimizes author style embeddings by integrating candidate space reduction with dimension selection strategies, enabling precise attribution via cosine similarity. Experimental results demonstrate that the LISA framework significantly outperforms existing baselines, confirming the critical role of high-quality stylistic representations in enhancing zero-shot authorship attribution performance.
📝 Abstract
Authorship Attribution (AA) requires capturing fine-grained stylistic characteristics, making it particularly challenging in zero-shot (ZS) settings where no task-specific supervision is available. In this work, we investigate the effect of author representations on ZS AA by evaluating a label-only prompting baseline together with three author representation strategies: representative writing samples, LLM-generated descriptions, and style embeddings (LISA). The first three approaches perform attribution using LLM prompting, while the embedding-based approach uses style embeddings with cosine similarity. We investigate the influence of prompt design and propose a two-stage embedding-based attribution framework that combines candidate space reduction with embedding-dimension selection. The results show that label-only ZS AA is ineffective, while incorporating author-specific representations consistently improves attribution performance. Among the evaluated approaches, the proposed two-stage LISA framework achieves the strongest overall performance, whereas LLM-generated style descriptions provide a substantially more compact representation of author style at the cost of some attribution performance. These findings demonstrate the importance of author representation in ZS AA, while indicating that current open-source LLMs remain insufficient for robust attribution without more effective representation learning.