Institution profile

National Defense Academy of Japan

Academic institutionasia · jp
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

May 27, 2026

This work addresses the challenge of visualizing the implicit periodic phase structures—such as stance and swing phases—embedded in deep reinforcement learning policies for locomotion control. To this end, the authors propose a feature augmentation mechanism that incorporates multi-step temporal information by extending state observations to include the current action, next state, and next action. Combined with a clustering approach that suppresses self-transitions and automatically determines the optimal number of clusters, this method enables the automatic extraction of latent phase structures from policy trajectories. Evaluated on MuJoCo’s Ant-v5, HalfCheetah-v5, and Walker2d-v5 environments, the approach significantly enhances the clarity of identified phases and the regularity of phase transition rules, outperforming existing methods.

0 citationsRead paper

Uncovering Latent Phase Structures and Branching Logic in Locomotion Policies: A Case Study on HalfCheetah

Mar 18, 2026

This work addresses the limited interpretability of deep reinforcement learning policies in locomotion control, despite their strong performance. Focusing on policies trained in the HalfCheetah-v5 environment, the study identifies semantically meaningful phases through clustering based on state similarity and transition consistency, revealing that neural networks can spontaneously develop human-interpretable periodic phase structures and branching logic without explicit supervision. To further elucidate these emergent behaviors, the authors employ Explainable Boosting Machines (EBM) to model each phase, explicitly uncovering the critical state features and action rules governing phase-specific control. This approach substantially enhances policy interpretability and offers a novel perspective for understanding the implicit motion planning strategies learned by autonomous agents.

0 citationsRead paper

Accuracy-Preserving CNN Pruning Method under Limited Data Availability

Nov 13, 2025

In data-scarce scenarios, conventional LRP-based CNN channel pruning suffers from substantial accuracy degradation and limited pruning ratios. To address this, we propose a fine-tuning-free, interpretability-driven pruning framework. Our method dynamically refines per-layer channel relevance scores by jointly leveraging structural priors from pre-trained models and an LRP-based importance re-evaluation mechanism—without requiring additional labeled data. This yields more robust channel importance estimation under low-data conditions. Experiments on ImageNet subsets and CIFAR benchmarks demonstrate that our approach achieves, on average, a 23.6% higher pruning ratio and reduces accuracy loss by 58.4% compared to state-of-the-art LRP-based pruning methods. It thus significantly overcomes the performance bottleneck of traditional LRP pruning in few-shot settings, delivering a practical, high-accuracy, high-compression solution suitable for resource-constrained edge deployment.

0 citationsRead paper
Recent publications

Latest Papers

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

May 27, 2026

This work addresses the challenge of visualizing the implicit periodic phase structures—such as stance and swing phases—embedded in deep reinforcement learning policies for locomotion control. To this end, the authors propose a feature augmentation mechanism that incorporates multi-step temporal information by extending state observations to include the current action, next state, and next action. Combined with a clustering approach that suppresses self-transitions and automatically determines the optimal number of clusters, this method enables the automatic extraction of latent phase structures from policy trajectories. Evaluated on MuJoCo’s Ant-v5, HalfCheetah-v5, and Walker2d-v5 environments, the approach significantly enhances the clarity of identified phases and the regularity of phase transition rules, outperforming existing methods.

0 citationsRead paper

Uncovering Latent Phase Structures and Branching Logic in Locomotion Policies: A Case Study on HalfCheetah

Mar 18, 2026

This work addresses the limited interpretability of deep reinforcement learning policies in locomotion control, despite their strong performance. Focusing on policies trained in the HalfCheetah-v5 environment, the study identifies semantically meaningful phases through clustering based on state similarity and transition consistency, revealing that neural networks can spontaneously develop human-interpretable periodic phase structures and branching logic without explicit supervision. To further elucidate these emergent behaviors, the authors employ Explainable Boosting Machines (EBM) to model each phase, explicitly uncovering the critical state features and action rules governing phase-specific control. This approach substantially enhances policy interpretability and offers a novel perspective for understanding the implicit motion planning strategies learned by autonomous agents.

0 citationsRead paper

Accuracy-Preserving CNN Pruning Method under Limited Data Availability

Nov 13, 2025

In data-scarce scenarios, conventional LRP-based CNN channel pruning suffers from substantial accuracy degradation and limited pruning ratios. To address this, we propose a fine-tuning-free, interpretability-driven pruning framework. Our method dynamically refines per-layer channel relevance scores by jointly leveraging structural priors from pre-trained models and an LRP-based importance re-evaluation mechanism—without requiring additional labeled data. This yields more robust channel importance estimation under low-data conditions. Experiments on ImageNet subsets and CIFAR benchmarks demonstrate that our approach achieves, on average, a 23.6% higher pruning ratio and reduces accuracy loss by 58.4% compared to state-of-the-art LRP-based pruning methods. It thus significantly overcomes the performance bottleneck of traditional LRP pruning in few-shot settings, delivering a practical, high-accuracy, high-compression solution suitable for resource-constrained edge deployment.

0 citationsRead paper