semantics-preserving code obfuscation

Designs, implements, and evaluates program transformations that change a program’s structure, identifiers, or representation while preserving its functional semantics and observable behavior. These techniques aim to conceal implementation details—by altering control and data flow, naming, or formatting—while optionally producing output that remains human-readable or is tuned to evade automated code-analysis tools.

semantics-preservingcodeobfuscation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.2
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Can LLMs Recover Program Semantics? A Systematic Evaluation with Symbolic Execution

Nov 24, 2025
RF
Rong Feng
🏛️ The Pennsylvania State University

Large language models (LLMs) struggle to recover the original semantics of programs obfuscated or optimized by compilers—e.g., via control-flow flattening, opaque predicates, or arithmetic/branch encoding—hindering program understanding, testing, and vulnerability analysis. To address this, we propose a symbolic-execution-augmented LLM semantic recovery method: leveraging KLEE to generate SMT constraints, path coverage statistics, and concrete test cases, we construct high-fidelity training data to fine-tune an LLM for joint modeling of syntactic structure and semantic constraints. This work is the first to systematically inject precise semantic information from symbolic execution into LLM training. Experiments on our custom obfuscation benchmark show that after fine-tuning, GPT-4.1-mini achieves a 37.2% improvement in compilation success rate and 89.6% behavioral equivalence—substantially outperforming state-of-the-art text-only or static-analysis baselines.

Assessing semantic preservation and compilation success across obfuscation transformationsEvaluating LLMs' ability to recover program semantics from obfuscated codeInvestigating symbolic execution-enhanced fine-tuning for program deobfuscation

This work proposes a novel approach to program transformation—such as compiler optimizations—by systematically integrating non-determinism and partially defined operations from functional logic programming. Leveraging the Curry language and its FlatCurry intermediate representation, the method expresses transformation rules in a concise, declarative style through pattern matching and non-deterministic computation. Traditional approaches often involve intricate manipulations of intermediate representations like abstract syntax trees, leading to implementations that are both cumbersome and error-prone. In contrast, the proposed technique enhances expressiveness and readability while maintaining practical feasibility. Experimental evaluation demonstrates that this approach not only preserves code clarity but also achieves competitive performance in real-world transformation tools, offering a compelling alternative for implementing reliable and maintainable program transformations.

abstract syntax treefunctional logic programmingintermediate representation

This study investigates the impact of code obfuscation on human program comprehension. Through output prediction experiments in Python and JavaScript, it systematically evaluates how multi-level obfuscation techniques—including identifier renaming, adversarial naming, and control flow modifications—affect comprehension accuracy and response time, as well as their interaction with programming experience. The findings reveal that obfuscation generally impairs comprehension efficiency, yet certain renaming strategies in Python unexpectedly enhance understanding. Obfuscation also shifts cognitive strategies from heuristic to more deliberative processing. Moderate thinking time correlates with optimal performance, and while programming experience confers benefits within a language, its transfer across languages is limited. Notably, the relationship between obfuscation strength and comprehension difficulty is non-monotonic and exhibits language-specific characteristics.

code obfuscationhuman reasoningoutput prediction

Same Same But Different: Preventing Refactoring Attacks on Software Plagiarism Detection

Oct 28, 2025
RM
Robin Maisch
🏛️ KASTEL – Institute of Information Security and Dependability | Karlsruhe Institute of Technology (KIT) | KTH Royal Institute of Technology

AI- and algorithm-driven semantic-preserving code obfuscation—such as variable renaming and control-flow restructuring—severely undermines the robustness of existing plagiarism detection systems in programming education. Method: We propose the first scalable detection framework integrating Code Property Graphs (CPGs) with graph transformation techniques. It constructs fine-grained CPGs via static analysis, formally models common refactoring operations, and employs invertible graph transformations to achieve semantic alignment and matching between obfuscated and original code. Contribution/Results: Our approach overcomes limitations of syntax- or shallow-semantic–based methods. Evaluated on a real-world student code dataset, it significantly improves detection accuracy for both AI-generated and manually refactored obfuscated code. Notably, it demonstrates superior generalizability and robustness against functionally equivalent structural modifications—e.g., those preserving program behavior while altering syntactic or control structures.

Detecting code plagiarism against refactoring obfuscation attacksEnhancing detectors using code property graphs and transformationsImproving resilience to structural modifications preserving behavior

ChangeGuard: Validating Code Changes via Pairwise Learning-Guided Execution

Oct 21, 2024
LG
Lars Gröninger
🏛️ University of Stuttgart

Detecting semantic-breaking modifications in non-functional changes—such as code refactoring and performance optimization—remains challenging due to their subtle behavioral impact. To address this, we propose a pairwise learning-guided execution framework that jointly models pre- and post-change behavioral discrepancies via dynamic execution monitoring, neural program embedding, pairwise contrastive learning, and mutation-driven input generation. Unlike conventional regression testing (which achieves only 7.6% recall), our approach is the first to integrate pairwise contrastive learning into the guided execution paradigm, significantly enhancing robustness and path coverage. Evaluated on 224 real-world, manually labeled code changes and three sets of automated transformations, our method achieves 77.1% precision and 69.5% recall—substantially outperforming baseline approaches—and successfully identifies unintended behavioral regressions introduced by mainstream automated refactoring tools.

Detect semantics-changing code modificationsImprove automated code transformation qualityValidate code behavior preservation

Latest Papers

What's happening recently
View more

This work addresses the vulnerability of Erlang programs to reverse engineering, decompilation, and recompilation attacks by proposing a multi-layered obfuscation scheme that jointly applies transformations at the source code, abstract syntax tree (AST), BEAM assembly, and bytecode levels. Leveraging the representational gap between Erlang’s high-level semantics and its low-level execution model, the approach introduces novel obfuscation paradigms based on opcode dependencies, encoded receive loops, and irregular control flow, further enhanced with dynamic module loading and self-modifying code techniques. The method operates fully within the constraints of the standard Erlang compiler, validator, loader, and virtual machine, thereby preserving compatibility while significantly strengthening resistance against both static and dynamic analysis, offering a practical and stealthy defense mechanism.

BEAM bytecodedecompilationErlang

Program analyses often lack robustness in the face of code changes. This work introduces, for the first time, a unified framework grounded in category theory that formalizes programs and their properties as categorical objects, capturing various forms of robustness—such as variable renaming and semantic refinement—via structure-preserving functors. Two implementation pathways are proposed: one lifts constructions from restricted computational models to general-purpose programs, while the other ensures stability in the composition of robust operators within algebraic program analyses. The framework not only uncovers common principles underlying loop summarization and termination analysis but also provides a theoretical foundation and predictability guarantees for developing program analyses that are more resilient to program transformations.

category theoryprogram analysisrobustness

Hot Scholars

JS

Jing Shao

Research Scientist, Shanghai AI Laboratory/Shanghai Jiao Tong University
Computer VisionMulti-Modal Large Language Model
CT

Cong Tian

Xidian University
Formal methodsProgram verificationSoftware engineering
AH

Amir Houmansadr

University of Massachusetts Amherst
Privacy-enhancing technologiesTrustworthy MLNetwork traffic analysis
AN

Ali Naseh

University of Massachusetts Amherst
Trustworthy Machine Learning