Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of non-smooth action oscillations that arise when deploying deep reinforcement learning policies on physical robots. To this end, we propose CATS, a method that introduces temporal regularization to constrain discrepancies between consecutive actions. We theoretically demonstrate that temporal penalties effectively bound spatial smoothness. Furthermore, a linearly increasing scheduling mechanism is designed to optimize the training process, balancing policy exploration with convergence stability. Experimental results indicate that CATS significantly suppresses action oscillations in both simulated and real-world environments while maintaining high task returns with minimal computational overhead. This work thus provides an efficient solution for the safe deployment of reinforcement learning on physical robotic systems.
📝 Abstract
Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoothness. Building on this finding, we propose Conditioning for Action using only Temporal Smoothness (CATS), which combines a temporal penalty with linear ramp-up. We highlight temporal regularization's ability to provide spatial smoothness while better preserving task performance than explicit spatial regularization. Through linear ramp-up, CATS allows the policy to learn rewarding behavior before progressively smoothing its actions, improving return preservation and both temporal and spatial smoothness. Experiments in both simulation and the real world show that CATS substantially reduces action oscillation without degrading task performance, with little computational overhead.
Problem

Research questions and friction points this paper is trying to address.

Deep Reinforcement Learning
Action Oscillation
Temporal Regularization
Spatial Smoothness
Physical Robot Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deep Reinforcement Learning
Temporal Regularization
Spatial Smoothness
Action Oscillation
CATS
🔎 Similar Papers
No similar papers found.
S
SungJae Ahn
College of Software, Kyung Hee University
J
Jeong Woon Lee
College of Software, Kyung Hee University
K
Kyoleen Kwak
College of Software, Kyung Hee University
Hyoseok Hwang
Hyoseok Hwang
Kyung Hee University
computer visionmachine learningrobotics