Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of fixed penalties in traditional reinforcement learning for greenhouse control, which struggle to constrain long-term climate violations. We formulate greenhouse regulation as a Constrained Markov Decision Process (CMDP) and propose a Lagrangian safe reinforcement learning framework based on RCPO-PPO. The core innovation lies in pioneering the integration of Kolmogorov-Arnold Networks (KANs) in place of standard MLPs to enhance nonlinear representation, combined with sinusoidal temporal features to capture diurnal cycles, thereby achieving decoupled optimization of economic returns and climate safety. Experimental results demonstrate that, compared to baseline methods, the proposed approach reduces cumulative climate violations by 18.65% while increasing lettuce profit by 2.91%, effectively balancing production efficiency with long-term risk mitigation.
📝 Abstract
Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.
Problem

Research questions and friction points this paper is trying to address.

Greenhouse climate control
Constrained Markov Decision Process
Safe reinforcement learning
Cumulative constraint violations
Nonlinear coupling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Markov Decision Process
Lagrangian Safe Reinforcement Learning
Kolmogorov-Arnold Networks
Greenhouse Climate Control
RCPO-PPO
H
Hangzun Liu
College of Informatics, Huazhong Agricultural University, Wuhan 430070, China
Y
Yuling Fan
College of Informatics, Huazhong Agricultural University, Wuhan 430070, China
F
Fang Tian
College of Informatics, Huazhong Agricultural University, Wuhan 430070, China; Yunnan Modern Agricultural Industry Research Institute Co., Ltd., Kunming, Yunnan 650000, China
Z
Zhilong Bie
College of Informatics, Huazhong Agricultural University, Wuhan 430070, China
Z
Zaiwen Feng
College of Informatics, Huazhong Agricultural University, Wuhan 430070, China; Yunnan Modern Agricultural Industry Research Institute Co., Ltd., Kunming, Yunnan 650000, China
Yongliang Qiao
Yongliang Qiao
Australian Institute for Machine Learning (AIML) ,The University of Adelaide
Smart agricultureCausalityArtificial intelligenceAgricultural robotsIntelligent perception