🤖 AI Summary
This study addresses the scalability challenge in networked Markov decision processes, where the state-action space grows exponentially with the number of agents and system dynamics are typically unknown, rendering conventional reinforcement learning impractical. To overcome this, the authors propose a Taylor expansion-based Q-function approximation framework. By leveraging the finite speed of information propagation, they prove that the expansion coefficients decay exponentially with graph distance, enabling dimensionality reduction through graph-local truncation. Integrating linear least-squares temporal difference (LSTD) with neural temporal difference parameterizations, this approach eliminates the reliance on known dynamic models inherent in traditional spectral analysis, achieving scalable model-free reinforcement learning. Extensive evaluations on three multi-agent control benchmarks demonstrate that the proposed method matches or surpasses existing baselines while efficiently scaling to large-scale graph structures.
📝 Abstract
In Networked Markov Decision Processes, transition dynamics are often unknown and the state--action space grows rapidly with the number of agents. In this setting, Taylor representations naturally approximate $Q$-functions, but a naive order-$n$ expansion over $N$ agents requires $\Theta(N^n)$ coefficients. We justify these expansions under smooth expected future local rewards with controlled derivatives. Under this condition, finite-speed information propagation and discounting imply that local-critic Taylor coefficients decay exponentially with the graph distance to the farthest agent involved. Discarding distant-agent coefficients and marginalizing then yield scalable local Taylor representations with a bound controlled by graph locality. Building on these representations, we propose a scalable model-free actor--critic algorithm, establishing finite-sample critic and near-stationarity guarantees for a linear LSTD critic. We then introduce a more expressive neural TD parameterization. Unlike prior constructive spectral methods, our approach covers settings without access to a known local dynamics map, such as hidden switched linear--quadratic regulation. Across three control benchmarks, our method matches or outperforms spectral baselines while scaling efficiently to large graphs.