SpikeCredit: Temporal Credit Carrier for Reinforcement Learning with Sparse Rewards

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge in sparse-reward reinforcement learning, where delayed outcomes struggle to guide intermediate decisions. To this end, it proposes SpikeCredit, a spiking neural network-based framework that formally defines a Temporal Credit Carrier (TCC). By leveraging spike dynamics as a substrate for credit retention, the method constructs a closed-loop mechanism comprising fast and slow read-write dual pathways, integrating self-motion feedback constraints with credit-target trajectory alignment to optimize policy learning. Evaluated across multiple sparse-reward MuJoCo tasks, the framework achieves up to a 1781% improvement in Last10 returns over baselines. Notably, on the Swimmer task, its performance even surpasses that of dense-reward baselines, effectively overcoming the long-horizon credit assignment bottleneck.
📝 Abstract
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Sparse Rewards
Credit Assignment
Spiking Neural Networks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spiking Neural Networks
Sparse Rewards
Temporal Credit Carriers
Credit Assignment
Reinforcement Learning
Y
Yingchao Yu
School of Information and Intelligent Science, Donghua University, Shanghai, China
Pengfei Sun
Pengfei Sun
Imperial College London
Delay LearningCognitive Modellingneuromorphic computingSpeech Processing
Wenxuan Pan
Wenxuan Pan
Institute of Automation, Chinese Academy of Sciences
Brain-inspired Intelligence
W
Wei Chen
School of Engineering, Westlake University, Hangzhou, China
Y
Yitian Hong
School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China
K
Kuangrong Hao
School of Information and Intelligent Science, Donghua University, Shanghai, China
Y
Yaochu Jin
School of Engineering, Westlake University, Hangzhou, China