Style, Not Self: Surface Cues Explain Zero-Shot Code Attribution by Large Language Models

๐Ÿ“… 2026-09-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether large language models can recognize their own generated code and whether observed self-preference stems from genuine self-awareness. Through zero-shot pairwise comparison experiments on the MBPP and HumanEval benchmarks, the authors propose a rule-based style normalization method that strips superficial features such as comments and type hints to conduct attribution analysis. The findings reveal that modelsโ€™ self-attribution capabilities are primarily driven by shallow stylistic cues, notably code length. Following style normalization, attribution accuracy degrades to chance level, demonstrating that self-preference does not originate from intrinsic cognition. Accordingly, this work recommends that future research report balanced accuracy to mitigate confounds introduced by stylistic biases.
๐Ÿ“ Abstract
If a language model can recognize code it wrote, it may favor that code as a judge, and instances of one model monitoring each other could collude. We test this zero-shot on current commercial models. Five LLMs generate solutions to MBPP, HumanEval, and DS-1000, seven more to MBPP, and models act as evaluators in four tasks: picking their own solution from a pair, judging whether a single solution is their own, identifying which of two solutions a named model wrote, and judging quality blind. In the single-solution task, balanced accuracy is 49-58% for all 15 model-benchmark combinations, while raw accuracy (38-67%) mostly reflects how readily a model claims authorship. In the pairwise task, accuracy across 14 evaluator-opponent combinations correlates at r=0.93 with how often the evaluator's solution is longer. Attribution to a named model succeeds on some pairs and is consistently inverted on others. A rule-based normalization that strips docstrings, comments, type hints, and local names preserves Pass@1 and leaves ten of twelve re-tested results at chance; the other two follow a length difference it leaves, although a trained classifier still separates most normalized pairs. Claude Haiku's self-preference also disappears. We recommend reporting balanced accuracy, heuristic baselines, and label consistency.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Code Attribution
Large Language Models
Self-Preference
Surface Cues
Model Collusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Code Attribution
Surface Cues
Rule-based Normalization
Balanced Accuracy
Self-Preference