Score
Designs, builds, and evaluates methods and tools that infer which actor created a given piece of content or artifact and that assign maintainer or contributor roles; this includes author attribution and procedures to attribute authorship. Work covers linking author identities across communication channels, extracting and comparing authorship signals across media, and measuring attribution reliability.
This survey systematically reviews authorship analysis (AA) research from 2015 to 2024, focusing on the core tasks of author attribution and author verification. To address these, we synthesize methodological advances across feature engineering (e.g., n-grams, stylometric features), classical machine learning (e.g., SVM, Random Forest), deep learning (e.g., RNNs, Transformers), and large language models (via fine-tuning and prompt engineering), organizing insights into a four-dimensional framework: *method–feature–dataset–challenge*. Our contribution is threefold: first, we provide the first comprehensive taxonomy of multi-paradigm AA approaches, clarifying their evolutionary trajectories and applicability boundaries; second, we explicitly identify critical research gaps—including low-resource language processing, multilingual adaptability, cross-domain generalization, and detection of AI-generated text; third, we offer a principled theoretical foundation and actionable guidelines for developing robust, multilingual, and interpretable AA systems.
This study addresses the critical challenge of source code authorship identification in software engineering, security, and digital forensics. It presents a systematic literature review of 47 studies published between 2012 and 2025, offering the first integrated framework that unifies behavioral biometrics and code stylometry to capture both stylistic and behavioral features of programmers. Employing a systematic mapping approach combined with content analysis and clustering, the review reveals that existing research predominantly focuses on closed-world authorship attribution tasks, heavily relies on limited benchmark datasets, and exhibits significant gaps in authorship verification, open-world scenarios, and reproducibility. This work provides a structured synthesis of the field and offers methodological guidance for future research directions.
The rise of large language models (LLMs) has intensified authorship attribution challenges—namely, distinguishing human-authored text from LLM-generated content and resolving ambiguous attribution in human-AI collaborative writing. To address this, we propose the first four-category authorship taxonomy for the LLM era: human-authored, LLM-generated, LLM-attributed, and human-AI collaborative. We systematically survey detection methodologies across four paradigms: statistical features, neural representations, attribution graphs, and prompt engineering—covering models including BERT, RoBERTa, and Llama, as well as watermarking and probability calibration techniques. Further, we introduce a unified evaluation framework balancing cross-domain generalizability and decision interpretability, and establish the field’s first dynamically updated resource repository (llm-authorship.github.io). Our work provides both theoretical foundations and a practical roadmap for enhancing detection accuracy and transparency in LLM-era authorship attribution.
This study addresses the persistent problem of authorship misconduct in scholarly publishing. Using a large-scale empirical analysis of 81,000 articles published in *PLOS ONE* (2018–2023), it systematically assesses the accuracy of author contributions via CRediT-based textual analysis of contribution statements, augmented by statistical modeling and cross-regional demographic validation. The study identifies that 9.14% of papers exhibit inappropriate authorship—impacting over 14,000 researchers—constituting the first large-scale, evidence-based quantification of such misconduct. Results reveal significant geographic clustering of authorship irregularities, particularly in Asia, Africa, and Italy, and demonstrate strong associations with authors’ regional backgrounds and professional affiliations—especially scholars affiliated with industry or non-profit organizations. By shifting authorship ethics governance from normative advocacy to a data-driven paradigm, this work provides critical empirical evidence to inform global research integrity policies and institutional oversight frameworks.
This work proposes a unified natural language processing framework to address key challenges in academic integrity, including plagiarism, content fabrication, and authorship verification. The framework integrates four core stylometric tasks: classification of human- versus machine-generated text, distinction between single- and multi-author documents, detection of authorship changes within multi-author texts, and identification of contributing authors in collaborative writing. The study introduces and publicly releases the first academic text dataset generated using Gemini under two distinct instruction settings—standard and strict—and systematically evaluates how prompting strategies affect detection performance. Experimental results demonstrate that texts produced under strict instructions are significantly more adversarial, thereby increasing the difficulty of accurate identification. The code and dataset are made openly available, establishing a new benchmark for research on academic integrity.
This study addresses the tension between efficiency gains from AI writing assistants and the potential erosion of authors’ psychological ownership over their text. To mitigate this trade-off, the work proposes five design patterns aimed at preserving authorial identity: on-demand activation, micro-suggestions, voice anchoring, audience scaffolding, and decision-point provenance. Through an online controlled experiment, two strategies—role-based coaching and style personalization—were evaluated. Results indicate that style personalization significantly enhances psychological ownership (+0.43) and increases AI content adoption by 5%, whereas role-based coaching fails to counteract the decline in ownership. Notably, cognitive load is reduced without compromising text quality. This research advances a novel paradigm for collaborative writing systems that effectively balances productivity with authorial agency and identity.
The widespread adoption of generative AI blurs the boundaries of users’ actual contributions in creative processes, often leading to misperceptions of authorship. This work introduces the novel concept of “authorship calibration”—defined as users’ accurate self-assessment of their genuine contribution in human-AI collaboration—and presents an empirical analysis based on the CoAuthor dataset. The study reveals that frequent AI users systematically overestimate their own input, whereas infrequent users exhibit more accurate calibration, thereby uncovering a link between AI usage intensity and metacognitive bias. These findings offer a new theoretical lens and empirical foundation for understanding how generative AI reshapes human perceptions of creative agency and authorship.
This study systematically investigates the generalization gap of existing code authorship attribution methods when applied beyond competitive programming contexts to real-world classroom assignments. While state-of-the-art approaches achieve strong performance on benchmark datasets such as Google Code Jam—attaining 70.7% Top-1 accuracy among 1,000 authors—their effectiveness collapses dramatically in educational settings, dropping to 0.2% and 0.06% on university course assignments, levels nearly equivalent to random guessing. Using pre-trained Transformer models like CodeBERT, we conduct comprehensive benchmarking across multiple sources, including Google Code Jam, Kaggle, and a newly curated dataset of student homework submissions. Our findings reveal a pervasive and previously underappreciated cross-domain performance cliff, highlighting significant practical limitations of current techniques in authentic educational scenarios.
This study addresses the lack of scientifically grounded evaluation methods for assessing author-style personalization in large language models (LLMs), a gap that renders conventional metrics inadequate for capturing true stylistic fidelity. To remedy this, the work introduces authorship verification theory into LLM evaluation and proposes an integrated framework combining the LUAR authorship verification model, a decoupled trait-matching LLM-based evaluator, and classical function-word stylometric analysis. Experiments on 1,000 generated texts from 50 authors reveal a significant “authorship gap”: all inference-time personalization methods score markedly lower (0.484–0.508) than the cross-author baseline (0.626). Moreover, near-zero correlations (|r| < 0.07) among the three evaluation dimensions demonstrate that theoretically ungrounded assessments risk drawing misleading conclusions. This work establishes a calibrated, theoretically informed benchmark with absolute interpretability for style personalization evaluation.
This study addresses the fundamental challenge posed by AI-generated code to the long-standing assumption in software engineering that authorship implies understanding. The authors demonstrate, for the first time, that AI-assisted programming systematically undermines the validity of authorship as a proxy for knowledge, thereby rendering traditional knowledge metrics—such as the truck factor—ineffective. By integrating principles from software engineering measurement theory, knowledge modeling, and logical reasoning, the work reinterprets the semantics of version control data in the context of AI collaboration. The research reveals that existing measures of knowledge concentration no longer reflect actual comprehension in AI-augmented development environments and advocates for a paradigm shift toward metrics grounded in verifiable evidence of understanding. It further identifies the construction of system-level understanding metrics as a critical open problem for the field.