steering vector distillation

Design and build methods that distill historical experiences or procedural knowledge into compact, task-specific steering vectors — using contrastive or related objectives — that can be indexed and retrieved at runtime to modulate model behavior without full fine‑tuning. Analyze pipelines for encoding, indexing, and applying these training‑free procedural codes (experience‑based embedding distillation, subliminal learning, contrastive procedural distillation) and audit their effectiveness as steering signals.

steeringvectordistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Moonshine: Distilling Game Content Generators into Steerable Generative Models

Aug 18, 2024
YN
Yuhe Nie
🏛️ New York University | Southern University of Science and Technology

Current PCGML approaches suffer from weak controllability and heavy reliance on limited real-world data, resulting in unpredictable outputs with constrained diversity and quality. To address this, we propose Text-to-game-Map (T2M), a novel paradigm that— for the first time—transfers procedural game map generation algorithms to text-conditioned generative models via neural knowledge distillation. Our method leverages large language models (LLMs) to automatically annotate synthetic training data and jointly trains a diffusion model with a lightweight Five-Dollar network, enabling natural-language-instruction-driven map generation. The resulting model matches the original procedural algorithm in diversity, geometric accuracy, and content fidelity, while supporting real-time, fine-grained semantic control. This significantly enhances controllability, interpretability, and practical utility of PCGML systems.

Content QualityControllable GenerationData Limitations

Large language models (LLMs) acting as agents lack persistent procedural memory, hindering their ability to reliably execute tasks over extended interactions. This work proposes a training-free, implicit activation-guidance framework that, for the first time, models procedural memory as guidance vectors in the model’s activation space. By distilling task-specific skills from historical experiences through contrastive learning, the approach directly activates relevant internal neural mechanisms without relying on explicit instructions. This design circumvents the semantic gap between symbolic instructions and executable actions, enabling tight coupling between memory and execution while remaining complementary to explicit methods such as retrieval-augmented generation (RAG). Experiments demonstrate that the proposed method achieves performance on par with explicit instruction-based approaches across four agent benchmarks; when combined with such methods, it significantly enhances robustness. Moreover, the guidance vectors exhibit structured task logic within the activation space.

activation steeringagent memorylarge language models

Improving Instruction-Following in Language Models through Activation Steering

Oct 15, 2024
AS
Alessandro Stolfo
🏛️ ETH Zürich | Microsoft Research

This work addresses the limited capability of large language models (LLMs) to adhere to fine-grained instruction constraints—such as formatting, length, and keyword requirements—and their poor generalization across zero-shot or cross-model settings. To this end, we propose activation steering: a lightweight, inference-time intervention that computes layer-wise neural activation differences between instruction-present and instruction-absent conditions, yielding interpretable, transferable, and composable instruction vectors. Crucially, no model fine-tuning is required. Our key contribution is the first formulation of instructions as cross-model-transferable activation-difference vectors, enabling vector composition (e.g.,叠加 multiple constraints) and foundation-model enhancement. Extensive evaluation across four mainstream LLMs demonstrates substantial improvements in instruction-following accuracy. The method supports constraint-aware generation without explicit instructions, concurrent multi-constraint control, and knowledge transfer from instruction-tuned models to base models.

Controlling output format, length, and word inclusion constraintsEnhancing instruction-following in language models via activation steeringTransferring steering vectors from tuned to base models

Latest Papers

What's happening recently
View more

This study addresses how metaphorical instructions can inadvertently induce large language models to transfer inefficient algorithmic patterns from the source domain during code generation, leading to degraded cross-task performance. The work presents the first systematic investigation into this metaphor-induced algorithmic bias, introducing the MASC framework to deliberately elicit and analyze the phenomenon through metaphor restructuring, behavioral evaluation, hidden state analysis, and prototype detection. Findings reveal that the detrimental influence of metaphors stems from deep programmatic patterns rather than superficial linguistic features and can be precisely identified: MASC achieves high accuracy in detecting metaphor-derived coding skills and their associated inefficient implementations, demonstrating that model hidden states shift toward prototypical representations of suboptimal behaviors.

algorithmic steeringcode generationlarge language models

This work addresses a key limitation of existing self-distillation methods—such as SDPO—that rely solely on turn-level rewards and overlook the rich procedural knowledge embedded across interaction trajectories. To remedy this, we propose Process Memory Distillation (PMD), which for the first time systematically organizes raw behaviors, self-reflective strategies, and cross-problem behavioral patterns from the model’s own trajectories into a three-layer reusable memory structure. A memory-guided self-teaching mechanism enables the joint co-evolution of policy and memory. Integrating online self-reflection, multi-level memory construction, and memory-conditioned self-supervised distillation, PMD achieves substantial performance gains on Qwen3-8B and OLMo3-Instruct-7B, with improvements of +3.8–5.5% on SCIKNOWEVAL and +7.9–13.6% on LIVECODEBENCH. Ablation studies confirm that freezing either the policy or memory component degrades performance by over 10%, validating the necessity of their协同 evolution.

cross-episode signalsmemory distillationprocedural memory

This work addresses the inefficiency in reinforcement learning when group trajectory rewards are all-or-nothing, which deprives the agent of informative relative signals. To overcome this, the authors propose SKALD, a novel framework that leverages abstract skills as dense supervision signals. SKALD employs online self-distillation within a single Qwen3-Base model by constructing dual views: a “question-only” student and a “skill-card-guided” teacher, enabling skill-aware knowledge injection into shared weights without requiring privileged inputs at test time. The method further incorporates an experience-gated mechanism and an annealed exponentially tilted objective to mitigate distribution shift and uninformative rewards. Evaluated across five mathematical benchmarks, SKALD achieves average@8 improvements of +2.46, +4.85, and +12.01 on 0.6B, 1.7B, and 4B models, respectively; notably, zero-variance distillation recovers 84.7% of the gain for the 1.7B model, substantially outperforming GRPO, FLOP-matched baselines, and context-based skill exposure approaches.

abstract skillsgroup-relative signalreinforcement learning

This work formally introduces the “instruction-to-skill” learning problem and proposes a closed-loop framework that transforms multimodal, heterogeneous procedural instructions from the web into structured, executable, and self-evolving skills for agents. Built upon a frozen vision-language model, the framework enables continuous optimization without human intervention or reference-based scoring by leveraging trajectory-driven skill revision and an analyzer-guided early stopping mechanism. The authors establish MMG2Skill-Bench, the first benchmark for this task, demonstrating consistent and substantial improvements over baselines across six vision-language models and diverse tasks, with macro-average performance gains of 12.8–25.3 percentage points. The proposed early stopping strategy further reduces the number of required attempts by 25%–53%.

agent-executable skillsguide-to-skill learningin-the-wild guides

Hot Scholars

JB

Jia-Bin Huang

Capital One Associate Professor at University of Maryland
Computational PhotographyComputer VisionMachine Learning
RH

Rongjie Huang

FAIR, Zhejiang University
Multimedia ComputingSpeechNatural Language Processing
ZZ

Zhou Zhao

Zhejiang University
Machine LearningData MiningMultimedia Computing
YW

Yongqi Wang

Zhejiang University
SpeechAudioDeep Learning
PT

Pavan Turaga

Geometric Media Lab, Arizona State University
Computer VisionMachine LearningGeometryTopology