strategy distillation

Designs methods and representations that extract and encode reusable strategies from demonstrations or policies into structured strategy descriptions and guidance signals. Builds and analyzes algorithms that inject those strategies into policy learning—e.g., strategy-guided policy optimization—supporting transfer of strategy knowledge (for example via forward-KL), evaluation of behavior with and without strategies, and instance-level adaptive weighting of guidance to produce autonomous and guided trajectories.

strategydistillation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Traditional trajectory imitation is confined to replicating specific solution steps and struggles to transfer general reasoning capabilities. This work proposes Strategy-Guided Policy Optimization (SGPO), which elevates instance-level imitation to strategy-level distillation by extracting structured strategy descriptions from a strong teacher model. SGPO generates paired trajectories—with and without strategy guidance—and integrates a selective forward KL objective, proximal constraint optimization, and an adaptive instance weighting mechanism to enable efficient and stable knowledge transfer. Evaluated on four mathematical reasoning benchmarks, SGPO substantially outperforms standard supervised fine-tuning and online policy reinforcement learning baselines, achieving an average improvement of 2.2 points on Qwen2.5-7B-Instruct, thereby demonstrating its effectiveness and scalability.

generalizationlanguage modelsreasoning distillation

This study addresses the lack of interpretability in AI strategies for complex games by proposing a novel method for generating executable programmatic policies. Inspired by human cognitive mechanisms, the approach integrates reinforcement learning with program synthesis to transform implicitly learned agent policies into executable programs based on action sequences. Experimental evaluations conducted in chess and grid-world environments demonstrate that the proposed method effectively extracts and generates efficient action sequences directly from data. Consequently, this approach significantly enhances the transparency and interpretability of AI decision-making processes. By bridging the gap between opaque policy representations and human-readable programmatic logic, this work establishes a new paradigm for game data analysis and explainable artificial intelligence.

Explainable RepresentationsGame-playing StrategiesReinforcement Learning

Learning Strategy Representation for Imitation Learning in Multi-Agent Games

Sep 28, 2024
SL
Shiqi Lei
🏛️ Chinese Academy of Sciences | Korea Advanced Institute of Science and Technology

In multi-agent imitation learning, offline datasets often contain heterogeneous policies, while existing methods rely on player identity annotations or strong prior assumptions. To address these limitations, this paper proposes STRIL—a framework that learns strategy-aware trajectory representations via contrastive learning and latent-space clustering, without requiring player identifiers or restrictive assumptions. STRIL introduces interpretable metrics—consistency and dominance—to quantitatively evaluate strategy quality, and integrates an end-to-end dynamic reweighting and pruning mechanism to identify and select dominant demonstration trajectories. The framework is plug-and-play, compatible with standard imitation learning algorithms such as Behavior Cloning (BC) and Generative Adversarial Imitation Learning (GAIL). Empirical evaluation on two-player Pong, Limit Texas Hold’em, and Connect Four demonstrates substantial performance gains over baselines; trajectory representation separability improves by over 32%, and dominant strategy trajectories are accurately identified.

Detects diverse strategies in multi-agent game trajectoriesFilters sub-optimal data to improve imitation learningLearns strategy representations without player identification

Sample-Efficient Behavior Cloning Using General Domain Knowledge

Jan 27, 2025
FZ
Feiyu Zhu
🏛️ Carnegie Mellon University

To address low sample efficiency and poor cross-environment generalization in behavioral cloning, this paper proposes the Knowledge-Guided Imitation Model (KIM). KIM automatically transforms informal, natural-language expert knowledge into semantically precise, differentiable structured policies via large language models, then jointly optimizes these policies with a minimal set of demonstration trajectories (only five). Crucially, KIM enables end-to-end encoding of informal domain knowledge into learnable policy structures—the first such approach—thereby realizing knowledge- and data-coordinated imitation learning. Evaluated on lunar landing and racing control tasks, KIM significantly outperforms knowledge-free baselines while demonstrating robustness to action noise. These results validate its high sample efficiency and strong generalization capability across diverse environments.

Behavioral CloningExpert Knowledge IntegrationLimited Examples

Reinforcement Learning via Implicit Imitation Guidance

Jun 09, 2025
PD
Perry Dong
🏛️ Stanford University

Offline demonstration-guided sample-efficient reinforcement learning suffers from performance degradation due to the misalignment between behavior cloning and reward optimization objectives in existing imitation learning approaches. Method: We propose Data-Guided Noise (DGN), the first method that leverages offline demonstrations solely to identify high-value action directions—fully decoupling exploration guidance from reward maximization. Within a policy gradient framework, DGN dynamically injects action-space noise informed by statistical analysis of demonstration trajectories, implicitly steering policy exploration toward high-return regions without explicit imitation. Contribution/Results: Evaluated on seven continuous-control benchmark tasks, DGN achieves 2–3× higher sample efficiency than state-of-the-art offline RL algorithms, while significantly improving policy generalization and asymptotic performance.

Avoiding imitation learning degradation of long-term performanceGuiding exploration via noise instead of behavior cloningSample efficient reinforcement learning with prior data

Latest Papers

What's happening recently
View more

This study addresses the lack of interpretability in reinforcement learning-based chess agents by extracting human-understandable tactical knowledge from black-box models. Methodologically, it innovatively integrates chess tactical priors to construct a symbolic sub-policy model and employs the PAL inductive logic programming system to extract board patterns. Furthermore, a novel divergence metric and computational evaluation scheme are proposed to validate the model's effectiveness. The results demonstrate that this approach successfully derives a set of tactically viable strategies whose move recommendations approximate those of novice human players. By maintaining decision interpretability, this work achieves an effective integration of symbolic AI and reinforcement learning.

Chess TacticsInductive Logic ProgrammingInterpretability

This work addresses the limited planning generalization of large language model (LLM) agents in unseen scenarios by proposing a dynamic policy learning framework that integrates generalized planning with hierarchical task decomposition. The approach automatically extracts and reuses parameterized policy components from successful executions to construct a composable policy library. Central to the method are hierarchical component learning (HCL-GP), semantic-driven policy retrieval, and a dynamic reuse mechanism that enables cross-task knowledge transfer. Evaluated on the AppWorld benchmark, the proposed method achieves task success rates of 98.2% on standard tasks and 97.8% on challenging ones—representing a 15.8 percentage point improvement over static composition. Notably, it elevates the success rate of open-source LLM agents from near zero to 62.5%, substantially enhancing their task generalization capabilities.

component generalizationgeneralized planninghierarchical task decomposition

Existing activation intervention methods struggle to effectively steer agent behavior even in simple decision-making tasks. This work formulates behavioral intervention as a reinforcement learning problem for the first time, constructing removable, composable, and reversible task vectors by accumulating policy gradients toward temporary behavioral objectives over a small number of trajectories. The resulting framework enables dynamic behavioral modulation and supports cross-task customization and composition of behaviors. Empirical validation demonstrates calibrated and reversible interventions in grid-world environments, flexible composition of tactical objectives in chess, and successful modification of team-specific behaviors in a football simulation setting, with effective generalization across diverse opponents.

activation steeringbehavioral controlinference-time intervention

This study addresses the challenge of balancing expert guidance with autonomous exploration in on-policy reinforcement learning, where existing methods often converge to suboptimal solutions due to fragmented optimization objectives. To overcome this, we propose an adaptive expert-guidance mechanism that treats the expert intervention weight as a learnable parameter, jointly optimized with the policy under a unified on-policy objective. Built upon Proximal Policy Optimization (PPO) and an alternating control framework, our approach enables automatic decay of expert influence without requiring auxiliary components or complex scheduling heuristics. Extensive evaluations across 34 tasks demonstrate significant improvements in sample efficiency over strong baselines, with notably low hyperparameter sensitivity. Crucially, the expert weight naturally diminishes to zero as performance improves, facilitating a smooth transition from reliance on expert demonstrations to independent policy execution, ultimately yielding policies that surpass the guiding expert.

Control BalancingExpert GuidanceOn-Policy Reinforcement Learning

Behavioral cloning often inherits unsafe or undesirable behavioral modes from demonstration data. To address this, this work proposes MoRE, a method that incorporates a brief “anti-cloning” phase to distill redirection signals—provided by a pattern classifier—into the policy weights, thereby guiding the policy toward desired behaviors without incurring additional inference overhead. MoRE is the first approach to eliminate undesirable modes without requiring intervention during inference, while preserving task performance, offering both efficiency and broad applicability. Compatible with architectures such as Diffusion Policy and Pi0.5 VLA, MoRE leverages classifier-guided distillation and a retention loss design, achieving an average 44-percentage-point improvement in deployment success across eight simulated and real-world tasks. Its performance approaches that of retraining baselines while maintaining original inference speed and task capabilities.

behavior cloninginference-time overheadpolicy adaptation

Hot Scholars

PY

Pin-Yu Chen

Principal Research Scientist, IBM Research AI; MIT-IBM Watson AI Lab; RPI-IBM AIRC
AI SafetyGenerative AITrustworthy Machine LearningAdversarial Machine Learning
TY

Tsung-Yi Ho

Chinese University of Hong Kong
Electronic Design AutomationMicrofluidicsTrustworthy Machine Learning
SP

Stjepan Picek

Faculty of Electrical Engineering and Computing, Croatia & Radboud University, The Netherlands
symmetric cryptographyAI Security and Privacyside-channel analysisBoolean functions
ZH

Zhiyuan He

The Chinese University of Hong Kong
Robust AIAdversarial Example
AN

Arun Narayanan

Senior Staff Research Scientist at Google Deepmind
Automatic Speech RecognitionSpeech AnalysisMachine LearningComputational Auditory Scene