Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6

๐Ÿ“… 2026-04-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether large language models retain poisoned identifiers in JavaScript deobfuscation even when correctly interpreting semantic meaning. Conducting 192 controlled experiments with Claude Opus 4.6 on two code prototypes, the authors employ task-specific prompt engineering, matched-pair design, and multi-step reasoning to trace identifier propagation and examine its relationship with semantic consistency. Results show that under baseline conditions, poisoned identifiers persist in 100% of outputs despite accurate semantic annotations; however, reformulating task instructions significantly reduces propagation rates to 0โ€“20%, demonstrating the efficacy of strategic prompting. The findings highlight string-table domain consistency as a critical factor, revealing that accurate semantic understanding does not inherently lead to identifier sanitization, thereby offering new insights for security-aware code generation.

Technology Category

Natural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)Cognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
๐Ÿ“ Abstract
When an LLM deobfuscates JavaScript, can poisoned identifier names in the string table survive into the model's reconstructed code, even when the model demonstrably understands the correct semantics? Using Claude Opus 4.6 across 192 inference runs on two code archetypes (force-directed graph simulation, A* pathfinding; 50 conditions, N=3-6), we found three consistent patterns: (1) Poisoned names persisted in every baseline run on both artifacts (physics: 8/8; pathfinding: 5/5). Matched controls showed this extends to terms with zero semantic fit when the string table does not form a coherent alternative domain. (2) Persistence coexisted with correct semantic commentary: in 15/17 runs the model wrote wrong variable names while correctly describing the actual operation in comments. (3) Task framing changed persistence: explicit verification prompts had no effect (12/12 across 4 variants), but reframing from "deobfuscate this" to "write a fresh implementation" reduced propagation from 100% to 0-20% on physics and to 0% on pathfinding, while preserving the checked algorithmic structure. Matched-control experiments showed zero-fit terms persist at the same rate when the replacement table lacks a coherent alternative-domain signal. Per-term variation in earlier domain-gradient experiments is confounded with domain-level coherence and recoverability. These observations are from two archetypes on one model family (Opus 4.6 primary; Haiku 4.5 spot-check). Broader generalization is needed
Problem

Research questions and friction points this paper is trying to address.

poisoned identifiers
LLM deobfuscation
code semantics
identifier persistence
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

prompt engineering
identifier poisoning
LLM deobfuscation
semantic coherence
code reconstruction
๐Ÿ”Ž Similar Papers