Mesh-RL: Coupled subgrid reinforcement learning

๐Ÿ“… 2026-06-24
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In large-scale or sparse-reward environments, reinforcement learning often suffers from slow value propagation due to the locality of temporal difference (TD) updates. This work proposes a finite element methodโ€“inspired spatial domain decomposition framework that partitions the state space into overlapping subdomains and enforces boundary-consistent TD updates, enabling efficient local learning while preserving global value consistency. By introducing, for the first time in reinforcement learning, the scientific computing principle of boundary-consistent domain decomposition, the method accelerates long-range credit assignment without modifying the reward function, Bellman operator, or incorporating explicit planning mechanisms. Evaluated on hazardous, dense grid worlds with diverse geometric structures, the framework significantly improves convergence speed, cumulative reward, and stability of algorithms such as Q-learning, SARSA, and Dyna-Q. Moreover, multi-resolution grids effectively mitigate premature convergence and enhance distant value propagation.
๐Ÿ“ Abstract
Reinforcement learning in large or sparse-reward environments suffers from slow temporal-difference reward propagation, as value information spreads only locally across the state space. We propose Mesh-RL, a spatial domain-decomposition framework inspired by the finite element method and domain decomposition theory, which partitions the environment into overlapping subgrids and enforces boundary-consistent temporal-difference updates. Such an approach enables localized learning while ensuring globally coherent value propagation. Unlike hierarchical or model-based approaches, Mesh-RL accelerates long-range credit assignment without modifying the reward function, Bellman operator, or introducing explicit planning mechanisms. We evaluate Mesh-RL on hazard-dense grid-world environments with varying geometries and mesh resolutions. Across Q-learning, SARSA, and Dyna-Q, Mesh-RL consistently improves convergence speed, cumulative reward, and learning stability. Higher mesh resolutions sustain exploration, prevent premature convergence, and substantially accelerate value propagation to distant states. While Dyna-Q already benefits from internal planning, it still achieves additional gains under structured decomposition. Overall, Mesh-RL introduces a principled spatial domain-decomposition mechanism for accelerating temporal-difference learning. Our framework bridges finite element method-inspired boundary-consistency techniques from scientific computing with reinforcement learning to improve sample efficiency in sparse-reward environments. We will release source code of the study.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
sparse-reward environments
temporal-difference learning
credit assignment
value propagation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mesh-RL
spatial domain decomposition
temporal-difference learning
boundary-consistent updates
sparse-reward environments
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
B
Behnam Gheshlaghi
Independent Researcher
B
Bahador Rashidi
Independent Researcher
S
Shahin Atakishiyev
University of Alberta