design input representations

Design and implement mappings that convert raw signals into model-ready inputs, specifying encoding formats such as symbolic versus continuous codes, entangled versus disentangled representations, and route- or pathway-specific input layouts. Build and evaluate alternative input encodings and parameter-sharing schemes to analyze how different representations affect information matching across input routes, code readability, and downstream model behavior.

designinputrepresentations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.22
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study investigates how input representations—specifically symbolic tokens, index-based “oracle” encodings, and entanglement-aware vectors—affect the binding capacity of miniature Transformers in compositional generalization. Models with 6–10K parameters are trained on fully enumerated finite factorized worlds, ensuring strict information-matching and eliminating sampling variance. The findings reveal that zero-shot binding performance falls below random chance across all input paths, with distinct failure modes: symbolic inputs lose answers during readout, index-based inputs suffer from incorrect binding, and entangled inputs inherit input readability limitations. Few-shot learning efficiency is jointly determined by parameter sharing and encoding readability inherent to each input path. These results challenge the assumption that clean oracle encodings are optimal, highlighting the critical role of input representation in compositional generalization.

compositional bindingfew-shot learninginput pathways

Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction

Jun 05, 2025
ZY
Zesheng Ye
🏛️ University of Melbourne | Southeast University | IBM Research

Efficient, lightweight downstream adaptation of large language models remains challenging due to fragmented methodologies and lack of unifying principles. Method: This paper proposes a unified framework—“neural reprogrammability”—modeling parameter-free adaptation paradigms—including model reprogramming, prompt tuning, and prompt instruction—as targeted manipulations of information flow at interfaces such as input, intermediate layers, or context. Contribution/Results: We introduce the first cross-modal, architecture-agnostic four-dimensional taxonomy (format, location, operator, output alignment), revealing intrinsic unity among in-context learning, chain-of-thought, and related methods. By systematically integrating existing interface perturbation techniques—including input perturbation, token insertion, and example injection—we empirically validate their generality across multimodal foundation models. Our framework establishes foundational principles and provides actionable guidelines for lightweight, controllable, and interpretable model adaptation.

Categorizing adaptation approaches across key dimensions systematicallyExploring neural network sensitivity to interface manipulationsUnifying model adaptation techniques for pre-trained foundation models

The absence of a unified symbolic and quantifiable structural description of neural network representation spaces impedes principled understanding of their generalization mechanisms. Method: We introduce the first scalable information-theoretic framework, featuring structural primitives for mapping architecture and an efficient vector-space entropy estimation algorithm—scalable to models ranging from millions to 12B parameters. Contribution/Results: Our framework enables unified characterization of representational structure evolution, its correlation with generalization performance, and structural commonalities across paradigms—including multi-agent reinforcement learning, sequence models, and large language models (LLMs). Empirically, we uncover a deep structural parallelism between linguistic constraints and neural generalization structure, and establish an interpretable causal chain: “design choices → emergent structure → observed performance.” This work provides a new paradigm for quantitative analysis and controllable design of neural representations.

Develop quantitative methods to analyze mapping structures in modelsLack unified notation for neural network representational spacesNeed methods to describe representation structure and generalization

Model-Driven Quantum Code Generation Using Large Language Models and Retrieval-Augmented Generation

Aug 27, 2025
NS
Nazanin Siavash
🏛️ University of Colorado Colorado Springs

Quantum and hybrid quantum-classical software development faces challenges including high heterogeneity costs across platforms and insufficient developer expertise. Method: This paper proposes a large language model (LLM) approach integrating Model-Driven Engineering (MDE) with Retrieval-Augmented Generation (RAG). It injects UML model instances as structured knowledge into the RAG framework to enable end-to-end, automated generation of executable Qiskit code from formal models. The method extends MDE-RAG-LLM synergy to code-to-code translation. Contribution/Results: Leveraging a semantic retrieval library built from GitHub-hosted open-source quantum code and optimized prompt engineering, our approach achieves a fourfold improvement in CodeBLEU score. It significantly enhances generated code accuracy, executability, and semantic consistency—establishing a novel paradigm for automated quantum software development.

Enhancing code accuracy with Retrieval-Augmented Generation pipelinesGenerating quantum code from UML models using LLMsReducing development costs for quantum-classical software systems

Latest Papers

What's happening recently
View more

Existing constrained decoding approaches treat schemas solely as structural constraints, overlooking the potential influence of their linguistic formulation on large language model behavior. This work reframes structured generation as a multi-channel instruction problem, demonstrating that subtle adjustments to schema key wording can implicitly convey instructions to guide model outputs—without altering prompts or model parameters. We systematically reveal, for the first time, that schema phrasing serves as an effective implicit instruction channel. Furthermore, we find significant differences across model families in their sensitivity to prompt-level versus schema-level instructions, with their interaction exhibiting non-additive effects. Experiments show that Qwen substantially benefits from schema-level instructions in mathematical reasoning tasks, whereas LLaMA relies more heavily on prompt-level guidance, and combining both channels does not necessarily yield cumulative performance gains.

constrained decodinginstruction channellarge language models

This study investigates the use of large language models (LLMs) to automatically translate neutral graph representations of fluid systems into high-quality, functionally correct code executable in mainstream simulation environments such as WNTR and Modelica. The authors systematically evaluate ten state-of-the-art LLMs combined with six prompting strategies across multiple benchmark scenarios, assessing generated code through software quality metrics and simulation fidelity. This work presents the first systematic comparison in the domain of fluid system modeling that examines how different LLMs and prompt engineering techniques influence both syntactic correctness and functional fidelity of generated simulation code, offering empirical guidance for model-driven code generation. Experimental results demonstrate that optimal configurations can produce syntactically valid code; however, a significant gap remains in achieving high simulation fidelity, highlighting key directions for future improvement.

code synthesisfluid systemslarge language models

This work addresses the challenge of emergent, poorly understood misalignments in fine-tuned language models that lead to unsafe code generation. The authors propose an interpretable activation-space framework that identifies actionable directions shared across architectures to detect and intervene on such misaligned behaviors. They demonstrate for the first time that internal misalignment directions exhibit causal specificity and manipulability, and uncover an asymmetric topological structure and a two-layer specificity mechanism underlying cross-architecture transfer. By integrating mean-difference directions, causal interventions, ridge regression mapping, and controlled safe-code experiments, they achieve 99.6% activation separation across four model families. Intra-model interventions reduce code leakage risk by 21–51 points, while cross-architecture transfer suppresses it by up to 46 points, albeit with limited specificity.

activation directionscode spillovercross-architecture transfer

This study investigates whether language models can effectively communicate intermediate reasoning during inference by transmitting hidden activations rather than textual tokens. Focusing on multi-hop reasoning tasks with Pythia models ranging from 160M to 410M parameters, the authors train linear mappings to align normalized hidden states between sender and receiver models and explore injecting the translated activations into the receiver either additively or via replacement. Despite achieving cosine similarities as high as 0.97—indicating strong representational alignment—neither injection strategy improves downstream performance: additive injection yields no significant gains, while replacement consistently degrades it. This work presents the first controlled demonstration of direct cross-model activation transfer and reveals the counterintuitive finding that representational alignment alone is insufficient to enable effective causal communication.

activation transfercross-model communicationhidden representations

This study addresses the challenge of natural language–driven simulation model discovery by systematically investigating the impact of data representation, Transformer-based embedding models, and reranking strategies on retrieval performance. By constructing multimodal model metadata and leveraging standard information retrieval metrics, the work presents the first quantitative evaluation of open-source embedding models for this task. Experimental results demonstrate that the proposed approach achieves strong performance in recall@5 and nDCG@5, with reranking substantially enhancing effectiveness on complex queries. These contributions establish the first benchmark framework for AI-enabled model reusability, composability, and interoperability in simulation model retrieval.

AI-driven retrievalmodel discoverymodel reuse

Hot Scholars

JL

Junyang Lin

Qwen Team, Alibaba Group & Peking University
Natural Language ProcessingCross-Modal Representation LearningPretraining
BH

Binyuan Hui

Qwen Team, Alibaba Group
Large Language ModelsCodeLLMsReasoningAgent
JY

Jiaxi Yang

PhD student, SIAT, CAS, China
Natural Language ProcessingLarge Language Model
LY

Luxin Yan

Huazhong University of Science and Technology
Computer VisionImage ProcessingDeep Learning
QZ

Qingfu Zhang

Chair Professor, FIEEE, City University of Hong Kong
evolutionary computationmultiobjective optimizationcomputational intelligence