🤖 AI Summary
This study addresses the lack of practical active sequential hypothesis testing methods and the difficulty of achieving decentralized pure exploration in multi-agent systems by proposing the ICMAPE framework. This approach pioneers the integration of Bayesian learning with deep reinforcement learning (TD3), reformulating fixed-confidence identification as a reward based on inference confidence. Through the coordination of a centralized neural inference network and decentralized local policies, ICMAPE dynamically determines when to terminate data collection. Evaluated on synthetic benchmarks and a Maryland nitrate monitoring task, the proposed method attains target precision with fewer exploration steps, significantly enhancing pure exploration efficiency in multi-agent settings.
📝 Abstract
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-specified models, while there is currently a gap for practical multi-agent methods that can perform active sequential testing. We fill this gap with ICMAPE, a Bayesian learning-based framework for decentralized multi-agent pure-exploration driven by inference objectives. ICMAPE converts the fixed-confidence identification objective into a reward derived from inference confidence, so that standard reinforcement learning machinery can be applied to decentralized pure exploration. It jointly learns a centralized neural inference network that estimates a posterior distribution over hypotheses from global trajectory data, and decentralized policies that select actions from local observation histories and learn when to stop collecting data once the target confidence is reached. On two synthetic benchmarks and a Maryland nitrate concentration monitoring task based on real-world data, ICMAPE-TD3 achieves target accuracy with fewer exploration steps.