🤖 AI Summary
This work addresses the challenge in social dilemmas where execution noise obscures whether an opponent’s defection stems from malicious intent or random error, often triggering excessive retaliation under conventional strategies. The authors formulate a partially observable Markov decision process (POMDP) in which the opponent’s intention is modeled as a latent state and observed actions are noisy emissions. Within an active inference framework, they jointly optimize epistemic (intention inference) and pragmatic (reward maximization) objectives. This approach explicitly disentangles intention from execution noise for the first time in social dilemmas, introduces a decomposable cost function, and theoretically links the critical noise threshold for cooperation collapse to fixed points of prior learning. Experiments in symmetric-noise repeated prisoner’s dilemmas show the strategy consistently outperforms conditional cooperators; however, when both players infer intentions under high noise, belief interdependence can precipitate synchronous cooperation breakdown.
📝 Abstract
In noisy social dilemmas, intended actions are stochastically corrupted before execution, so an observed defection may reflect hostile intent or action error. Standard Markov Decision Process (MDP) formulations treat executed actions as states, structurally precluding this distinction and causing systematic over-retaliation. We introduce a Partially Observable MDP (POMDP) formulation encoding opponent intentions as latent states and executed actions as noisy observations, solved within the active inference (AIF) framework with a cost function that decomposes into epistemic and pragmatic components that jointly address inferring current intent and learning how intent evolves. In the Iterated Prisoner's Dilemma with symmetric noise, we derive a critical noise threshold governing cooperation collapse, connecting it to a fixed-point condition on learned priors. Experiments reveal that the value of intention inference is context-dependent: the POMDP provides consistent advantages against conditionally cooperative opponents, but mutual intention inference under sufficient noise produces correlated belief-driven collapse. The advantage is specific to games where intent attribution is decision-relevant.