Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses multi-agent safety and resource depletion in the tragedy of the commons by formulating renewable commons as Markov games with explicit conservation constraints. It proposes a non-stationary Lagrangian framework that integrates constrained IPPO/MAPPO algorithms to construct policy sequences. Furthermore, this work introduces the concepts of average epoch solutions and reward-agnostic feasibility certificates to quantify price dispersion terms arising from selfish agent deviations. Validated on the Gordon-Schaefer fishery model, the approach demonstrates how depletion budgets influence population preservation. Ultimately, it provides theoretical guarantees for constrained policy sequences derived from unconstrained solutions, enabling sustainable resource harvesting without collapse.
📝 Abstract
The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce average-epoch solution concepts for reset episodes with discounted rewards and terminal costs. We prove a reward-independent feasibility certificate, cooperative feasibility and approximate optimality against feasible policy mixtures, and an extension to unbiased sampled costs. For self-interested agents, a constrained Nash certificate quantifies the price-dispersion term introduced by deviations that redistribute budget across epochs. Under the stated assumptions on solver accuracy and multiplier updates, these results give constrained policy-sequence guarantees using solutions of unconstrained problems. Experiments with constrained IPPO and MAPPO in a Gordon-Schaefer fishery examine how depletion budgets shape stock retention, harvest rewards, and price adaptation.
Problem

Research questions and friction points this paper is trying to address.

Tragedy of the Commons
Multi-agent Safety
Constrained Markov Game
Resource Depletion
Regenerative Commons
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constrained Markov Game
Nonstationary Lagrangian Framework
Average-Epoch Solution
Multi-Agent Reinforcement Learning
Feasibility Certificate
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.