🤖 AI Summary
This study addresses the trade-off between detection rate and functional correctness in LLM code watermarking, as well as false positives caused by pattern repetition. To this end, it proposes a robust code attribution framework based on stylistic rules. Methodologically, a style encoding mechanism driven by secret keys and structural context is designed, integrated with human code style probability calibration and context-aware evidence aggregation to effectively suppress evidence inflation and reduce false alarms. Experimental results on the CodeContests benchmark demonstrate that the proposed approach improves TPR@FPR5% by 12.44% over baselines. Furthermore, under four categories of non-LLM editing attacks, performance degrades by only 0.94%, indicating both high detection accuracy and strong adversarial robustness.
📝 Abstract
Code watermarking supports provenance tracking for code generated by LLMs. Modifying token selection to embed watermarks as an LLM generates code can create a trade-off between detectability and functional correctness. Other methods instead watermark completed code using predefined transformations or trained neural models. Recurring patterns can make watermark choices predictable across programs, while treating patterns common in unwatermarked code as watermark evidence can cause false detections. We therefore introduce SEW, which embeds and detects watermarks in already generated code through three components: (i) code style rules collected from style guides and transformation rules, with style choices determined by a secret key and each program's structural context; (ii) style-preference calibration, which evaluates watermark evidence using style probabilities estimated from human-written code; and (iii) context-aware style aggregation, which combines evidence from structurally matching locations assigned the same code style choice, preventing repeated applications of that choice from inflating watermark evidence. On CodeContests across three LLMs and three programming languages, SEW achieves a mean relative improvement of 12.44% in TPR@FPR5% over the baselines and is robust to four non-LLM code-editing attacks, with only a 0.94% mean relative decrease. Our code is available at https://github.com/suhanmen/SEW.