🤖 AI Summary
AI- and algorithm-driven semantic-preserving code obfuscation—such as variable renaming and control-flow restructuring—severely undermines the robustness of existing plagiarism detection systems in programming education.
Method: We propose the first scalable detection framework integrating Code Property Graphs (CPGs) with graph transformation techniques. It constructs fine-grained CPGs via static analysis, formally models common refactoring operations, and employs invertible graph transformations to achieve semantic alignment and matching between obfuscated and original code.
Contribution/Results: Our approach overcomes limitations of syntax- or shallow-semantic–based methods. Evaluated on a real-world student code dataset, it significantly improves detection accuracy for both AI-generated and manually refactored obfuscated code. Notably, it demonstrates superior generalizability and robustness against functionally equivalent structural modifications—e.g., those preserving program behavior while altering syntactic or control structures.
📝 Abstract
Plagiarism detection in programming education faces growing challenges due to increasingly sophisticated obfuscation techniques, particularly automated refactoring-based attacks. While code plagiarism detection systems used in education practice are resilient against basic obfuscation, they struggle against structural modifications that preserve program behavior, especially caused by refactoring-based obfuscation. This paper presents a novel and extensible framework that enhances state-of-the-art detectors by leveraging code property graphs and graph transformations to counteract refactoring-based obfuscation. Our comprehensive evaluation of real-world student submissions, obfuscated using both algorithmic and AI-based obfuscation attacks, demonstrates a significant improvement in detecting plagiarized code.