🤖 AI Summary
This study systematically investigates the generalization gap of existing code authorship attribution methods when applied beyond competitive programming contexts to real-world classroom assignments. While state-of-the-art approaches achieve strong performance on benchmark datasets such as Google Code Jam—attaining 70.7% Top-1 accuracy among 1,000 authors—their effectiveness collapses dramatically in educational settings, dropping to 0.2% and 0.06% on university course assignments, levels nearly equivalent to random guessing. Using pre-trained Transformer models like CodeBERT, we conduct comprehensive benchmarking across multiple sources, including Google Code Jam, Kaggle, and a newly curated dataset of student homework submissions. Our findings reveal a pervasive and previously underappreciated cross-domain performance cliff, highlighting significant practical limitations of current techniques in authentic educational scenarios.
📝 Abstract
Source code authorship attribution aims to identify the author of a program fragment from its writing style. We fine-tune the pre-trained transformer CodeBERT on three sources of data: publicly available Google Code Jam (GCJ) submissions from an open Kaggle repository, a curated GCJ archive, and institutional coursework datasets collected at a technical university. On multi-round GCJ data, CodeBERT reaches 92.6% Top-1 accuracy for 10 authors and retains 70.7% Top-1 (88.2% Top-10) for 1000 authors. On the examined coursework datasets, the same pipeline performs at or below the corresponding chance baselines: 0.2% Top-1 on a closed-assignment dataset of 690 authors and 0.06% Top-1 on open-ended assignments evaluated over 812 authors. A companion multi-model benchmark provides consistent cross-model evidence across additional dataset configurations, indicating that the observed performance gap is not specific to CodeBERT and persists across the evaluated model families. We analyze dataset and task properties that plausibly explain this gap and argue that GCJ-based benchmarks overestimate the practical applicability of authorship attribution in educational settings unless they are validated on the target coursework context.