program representation learning

Designs and trains models that produce learned representations (e.g., vector embeddings) of programs or program components such as methods and attributes, optimizing those embeddings to encode semantic responsibilities, concerns, or other program-level semantics. Builds and evaluates the representation models and similarity or evaluation metrics used to analyze, compare, and reason about program elements.

programrepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a novel approach to program analysis and optimization leveraging large language models (LLMs). Addressing the challenge of effectively integrating source code and intermediate representation (IR) information—a limitation in existing methods—it introduces LLMCompiler, pre-trained on IR, and employs a chunked embedding and aggregation strategy to produce unified program-level embeddings. By innovatively unifying the semantics of source code and IR, the method achieves a 1.54% error rate on algorithm classification, representing a 12% improvement over the current state of the art. It also attains competitive accuracy in heterogeneous device mapping, significantly advancing the application of LLMs in program understanding and optimization.

code optimizationintermediate representationLarge Language Models

Behavioral Embeddings of Programs: A Quasi-Dynamic Approach for Optimization Prediction

Oct 15, 2025
HP
Haolin Pan
🏛️ University of Chinese Academy of Sciences | Institute of Software Chinese Academy of Sciences

Compiler optimization faces a challenge in program embedding: static representations lack behavioral sensitivity, while dynamic ones incur high overhead and suffer from irreproducibility—hindering accuracy and scalability. Method: This paper proposes a quasi-dynamic program representation framework. Its core innovation is the “program behavior spectrum,” captured via lightweight optimization-sequence probing to model how program behavior evolves under compiler transformations. It further applies product quantization (PQ) to discretize continuous behavioral responses into subword-like units and introduces a multi-task PQ-BERT model that jointly learns behavioral semantics under syntactic structure constraints. Results: On best-optimization-path prediction and -Oz benefit prediction, the method significantly outperforms state-of-the-art static approaches, achieving higher prediction accuracy and stronger cross-program generalization.

Creating quasi-dynamic program representations for compiler optimization predictionLearning behavioral embeddings using compositional quantization and transformersModeling program optimization sensitivity through static feature changes

The selection of source code representations for machine learning–driven cybersecurity tasks lacks systematic, evidence-based guidance. Method: We conduct the first comprehensive empirical study, constructing a multidimensional evaluation framework that integrates diverse code representations—including abstract syntax trees (ASTs), tokens, control-flow graphs (CFGs), and program dependence graphs (PDGs)—with models spanning SVMs, RNNs, GNNs, and Transformers, across vulnerability detection, malware classification, and other security tasks in C, Python, and Java. Contribution/Results: We introduce a four-dimensional mapping taxonomy—“representation–task–language–model”—and quantitatively demonstrate that graph-based representations (especially ASTs) achieve superior popularity and effectiveness; C-language programs and vulnerability detection dominate current research focus; and sequence models and SVMs remain the most widely adopted. Our analysis identifies critical representation-task-language-model alignment principles and exposes key research gaps, providing both theoretical foundations and practical guidelines for security-oriented code representation design.

Analyzing representation impact on model learning and feature extractionIdentifying popular representations, models, tasks, and languages in current researchSurveying source code representations for ML-based cybersecurity tasks

Learning Program Behavioral Models from Synthesized Input-Output Pairs

Jul 11, 2024
TM
Tural Mammadov
🏛️ CISPA Helmholtz Center for Information Security | Saarland University

This work addresses black-box program behavior modeling by proposing a reversible, differentiable, and constraint-aware, grammar-driven neural modeling framework. Methodologically, it generates input-output (I/O) pairs from formal grammars of input and output languages, and employs a lightweight (<6.3M-parameter) sequence-to-sequence model to cast program I/O mapping as a bidirectional neural machine translation task—enabling both forward prediction and backward inference—while supporting fine-grained behavioral constraints and fault- or coverage-guided input synthesis. Its key contribution is the first end-to-end joint modeling of reversibility, differentiability, and syntactic consistency in program behavior models. Evaluated on structured tasks such as Markdown and HTML generation, the framework achieves 95.4% accuracy and a BLEU score of 0.98±0.04, significantly outperforming existing irreversible or syntax-agnostic approaches.

Assists in program understanding and maintenance through synthesis.Learns program behavior models from input-output pairs.Predicts outputs and inputs using neural machine translation.

Can Large Language Models Understand Symbolic Graphics Programs?

Aug 15, 2024
ZQ
Zeju Qiu
🏛️ Max Planck Institute for Intelligent Systems | University of Cambridge | MIT

This work investigates the capacity of large language models (LLMs) to perform spatial semantic understanding and cross-modal reasoning solely from symbolic graphical programs—such as curve parameters, stroke sequences, and local curvature—without visual encoders. To this end, we introduce the first benchmark for symbolic graphical program-based visual understanding, comprising three tasks: program generation, semantic question answering, and zero-visual-input cross-modal reasoning. We propose Symbolic Instruction Tuning (SIT), a novel fine-tuning paradigm that explicitly enhances LLMs’ spatial reasoning capabilities using synthetic symbolic graphical instruction data. Experimental results demonstrate that strong reasoning-oriented LLMs achieve superior performance; SIT substantially improves accuracy on symbolic graphical understanding tasks and—unexpectedly—generalizes to multiple general-purpose reasoning benchmarks (e.g., GSM8K, MMLU), indicating that symbolic spatial representations can strengthen foundational reasoning abilities.

Assessing LLMs' spatial-semantic reasoning via symbolic graphics programsImproving LLMs' reasoning with Symbolic Instruction Tuning (SIT)Testing LLMs' ability to imagine visuals from symbolic descriptions

Latest Papers

What's happening recently
View more

This study addresses the lack of effective code embedding methods for visual programming languages such as Scratch by systematically evaluating four large language models combined with five embedding strategies on textual representations of Scratch programs. The authors construct datasets for token prediction and program functionality classification tasks, demonstrating that embedding models trained on large-scale Scratch data effectively integrate structural and semantic information. Notably, these embeddings accurately predict the functional correctness of student programs without requiring task-specific fine-tuning. The findings offer a transferable and efficient embedding framework for classroom-level learning analytics, thereby filling a critical gap in representation learning for visual programming education.

block-based programmingcode embeddingsembedding evaluation

This work addresses a critical limitation in existing binary code representation learning methods, which typically overlook instruction-level alignment information and thus fail to effectively leverage fine-grained supervisory signals from compiler debug information. To overcome this, the paper introduces the first approach that explicitly models instruction alignment as an auxiliary training objective. By employing multi-task learning, the method jointly optimizes function-level embeddings and instruction alignment, using debug information to construct precise alignment supervision signals. Experimental results demonstrate that this approach significantly improves accuracy in binary code similarity retrieval, enhances the model’s discriminative power and semantic understanding, and reveals a strong correlation between instruction alignment and the quality of function representations. Consequently, it establishes a more interpretable and precise framework for binary code representation learning.

binary code representation learningcompiler debug informationfine-grained correspondence

This study addresses the limitation of text embeddings in capturing structural similarities and varying levels of abstraction among creative ideas. To overcome this, we propose a structured decomposition-based evaluation framework that deconstructs ideas into core components—such as purpose and mechanism—and constructs multi-layered concept graphs. By introducing shared structural representations and set-level mechanism coverage metrics, the framework enables component-wise overlap analysis to precisely quantify distinctions between core mechanisms and implementation details. Experimental results demonstrate that our approach improves alignment with expert judgments by 31%, significantly enhancing the evaluation of both similarity and diversity in large-scale ideation tasks.

creativity evaluationidea diversityidea similarity

This work addresses critical limitations in current programming education approaches that rely on closed-source large language models (LLMs) to simulate student behavior, including privacy risks, high costs, and model dependency. The authors propose a novel method that first reformulates student programming process logs into alternating sequences of “code submissions” and “environmental feedback,” structured as dialogues. Leveraging this representation, they perform supervised fine-tuning and preference optimization on open-source LLMs such as Qwen-4B and Qwen-8B to better model authentic student debugging behaviors. Experimental results demonstrate that this approach significantly outperforms baseline methods—both code-only models and prompt-driven large models—in terms of functional alignment and code similarity, thereby substantially improving the fidelity of reconstructed student debugging trajectories.

debugging behaviorlanguage modelsopen-weight models

Hot Scholars

MK

Marta Kryven

Assistant professor
cognitive scienceartificial intelligencerepresentation learningprogram induction
ZS

Zeyu Sun

Institute of Software, Chinese Academy of Sciences
Software EngineeringNatural Language Processing
JM

Jiayuan Mao

MIT CSAIL
Artificial IntelligenceRoboticsComputer VisionNatural Language Processing
BH

Bowei He

City University of Hong Kong, MBZUAI
Data MiningLanguage ModelGenAI4ScienceAgentic AI