Source Code Authorship Attribution Does Not Generalize from Competitions to Classrooms

📅 2026-07-31
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically investigates the generalization gap of existing code authorship attribution methods when applied beyond competitive programming contexts to real-world classroom assignments. While state-of-the-art approaches achieve strong performance on benchmark datasets such as Google Code Jam—attaining 70.7% Top-1 accuracy among 1,000 authors—their effectiveness collapses dramatically in educational settings, dropping to 0.2% and 0.06% on university course assignments, levels nearly equivalent to random guessing. Using pre-trained Transformer models like CodeBERT, we conduct comprehensive benchmarking across multiple sources, including Google Code Jam, Kaggle, and a newly curated dataset of student homework submissions. Our findings reveal a pervasive and previously underappreciated cross-domain performance cliff, highlighting significant practical limitations of current techniques in authentic educational scenarios.
📝 Abstract
Source code authorship attribution aims to identify the author of a program fragment from its writing style. We fine-tune the pre-trained transformer CodeBERT on three sources of data: publicly available Google Code Jam (GCJ) submissions from an open Kaggle repository, a curated GCJ archive, and institutional coursework datasets collected at a technical university. On multi-round GCJ data, CodeBERT reaches 92.6% Top-1 accuracy for 10 authors and retains 70.7% Top-1 (88.2% Top-10) for 1000 authors. On the examined coursework datasets, the same pipeline performs at or below the corresponding chance baselines: 0.2% Top-1 on a closed-assignment dataset of 690 authors and 0.06% Top-1 on open-ended assignments evaluated over 812 authors. A companion multi-model benchmark provides consistent cross-model evidence across additional dataset configurations, indicating that the observed performance gap is not specific to CodeBERT and persists across the evaluated model families. We analyze dataset and task properties that plausibly explain this gap and argue that GCJ-based benchmarks overestimate the practical applicability of authorship attribution in educational settings unless they are validated on the target coursework context.
Problem

Research questions and friction points this paper is trying to address.

source code authorship attribution
generalization gap
educational settings
programming competitions
code stylometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

source code authorship attribution
CodeBERT
generalization gap
educational datasets
empirical evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Serhii Yemets
Technical University of Košice, Letná 9, 042 00 Košice, Slovakia
M
Marek Horváth
Technical University of Košice, Letná 9, 042 00 Košice, Slovakia