Investigating Efficacy of Perplexity in Detecting LLM-Generated Code

📅 2024-12-21
🏛️ arXiv.org
📈 Citations: 3
✨ Influential: 1
📄 PDF
🤖 AI Summary
Perplexity is widely adopted for detecting large language model–generated code (LLMgCode), yet its practical effectiveness—particularly regarding accuracy, generalizability, and efficiency—remains inadequately characterized. Method: We conduct a systematic empirical evaluation of perplexity across a multilingual benchmark comprising over 24,000 code snippets, assessing detection accuracy, cross-language and cross-task generalization, and inference latency. Contribution/Results: We find that perplexity achieves the strongest generalization—outperforming alternatives across languages and programming tasks—and offers intrinsic interpretability. However, its average detection accuracy remains low (<65%), inference is 3–8× slower than feature-engineering–based methods, and performance degrades markedly on high-level languages (e.g., Python, JavaScript) versus lower-level ones (e.g., C, Java). Our analysis delineates the fundamental trade-offs among generalization, accuracy, and efficiency, establishing a three-dimensional evaluation framework that informs both theoretical understanding and practical selection of LLMgCode detection techniques.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Knowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large language model-generated code (LLMgCode) has become increasingly prevalent in software development. Many studies report that LLMgCode has more quality and security issues than human-authored code (HaCode). It is common for LLMgCode to mix with HaCode in a code change, while the change is signed by only human developers, without being carefully checked. Many automated methods have been proposed to detect LLMgCode from HaCode, in which the perplexity-based method (PERPLEXITY for short) is the state-of-the-art method. However, the efficacy evaluation of PERPLEXITY has focused on the detection accuracy. In this article, we are interested in whether PERPLEXITY is good enough in a wider range of realistic evaluation settings. To this end, we devise a large-scale dataset that includes 11,664 HaCode snippets and 13,164 LLMgCode snippets, and based on that, we carry out a family of experiments to compare PERPLEXITY against feature-based and pre-training-based methods from three perspectives: (1) detection accuracy in terms of programming language, degree of difficulty, and scale of solution, (2) generalization capability, and (3) inference efficiency. The experimental results show that PERPLEXITY has the best generalization capability while it has low accuracy and efficiency in most cases. Based on the experimental results and detection mechanism of PERPLEXITY, we discuss implications into both the strengths and limitations of PERPLEXITY, e.g., PERPLEXITY is unsuitable for high-level programming languages while it has good interpretability. As the first large-scale investigation on detecting LLMgCode from HaCode, this article provides a wide range of evidence for future improvement.
Problem

Research questions and friction points this paper is trying to address.

Evaluating perplexity method for detecting LLM-generated code
Comparing detection accuracy, speed, and generalization of methods
Assessing suitability of perplexity for high-level programming languages
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluates perplexity method for LLM code detection
Compares accuracy, speed, generalization capabilities
Identifies strengths and limitations of PERPLEXITY
🔎 Similar Papers
Nanjing University | University of Zürich | University of Malaya
J
Jinwei Xu
State Key Laboratory of Novel Software Technology, Software Institute, Nanjing University, Nanjing 210008, China
H
He Zhang
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
Y
Yanjin Yang
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
Z
Zeru Cheng
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
J
Jun Lyu
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
B
Bohan Liu
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
X
Xin Zhou
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
Lanxin Yang
Lanxin Yang
State Key Laboratory of Novel software technology, Software Institute, Nanjing University, Nanjing 210008, China
Alberto Bacchelli
Alberto Bacchelli
Associate Professor, Head of ZEST @ University of Zurich
empirical software engineeringcode review
Y
Yin Kia Chiam
Department of Software Engineering, Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur 50603, Malaysia
T
Thiam Kian Chiew
Department of Software Engineering, Faculty of Computer Science and Information Technology, University of Malaya, Kuala Lumpur 50603, Malaysia