Institution profile

University of Latvia

Academic institutioneurope · lv
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

Interpretable Policy Distillation for Power Grid Topology Control

May 30, 2026

This work addresses the challenges of high inference cost, deployment difficulty, and opaque decision-making in deep reinforcement learning for power grid topology control. The authors propose a stress-focused data collection strategy to train a Proximal Policy Optimization (PPO) teacher model and, for the first time, distill it into interpretable, lightweight agents—specifically decision trees and random forests—targeting high-load critical states. The distilled models not only surpass the original PPO policy in average reward and survival duration while significantly reducing inference overhead, but also maintain highly consistent action outputs, enabling human auditability. Furthermore, the study reveals fundamental differences in feature dependencies between neural policies and tree-based models, achieving a balanced trade-off among performance, real-time responsiveness, and interpretability.

0 citationsRead paper
Recent publications

Latest Papers

Interpretable Policy Distillation for Power Grid Topology Control

May 30, 2026

This work addresses the challenges of high inference cost, deployment difficulty, and opaque decision-making in deep reinforcement learning for power grid topology control. The authors propose a stress-focused data collection strategy to train a Proximal Policy Optimization (PPO) teacher model and, for the first time, distill it into interpretable, lightweight agents—specifically decision trees and random forests—targeting high-load critical states. The distilled models not only surpass the original PPO policy in average reward and survival duration while significantly reducing inference overhead, but also maintain highly consistent action outputs, enabling human auditability. Furthermore, the study reveals fundamental differences in feature dependencies between neural policies and tree-based models, achieving a balanced trade-off among performance, real-time responsiveness, and interpretability.

0 citationsRead paper