Bounded Reasoning: Cognitive Hierarchy in Human-versus-AI Cyber Defense

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that single average scores in cybersecurity defense obscure behavioral discrepancies between humans and AI arising from distinct adaptation mechanisms. To investigate this, a sequential game model is constructed based on attack graphs, incorporating cognitive hierarchy theory to characterize information structures. By combining DQN and CHT-DQN reinforcement learning agents with Mechanical Turk crowdsourcing experiments, the evolution of human versus machine defense strategies under reward feedback is systematically compared. The findings reveal asymmetric adaptive behaviors in humans driven by prospect theory, demonstrating that evaluations relying solely on average scores overlook critical characteristics. Furthermore, the results indicate that human defenders outperform automated baselines in later defense stages, highlighting the necessity of nuanced evaluation metrics beyond aggregate performance measures in human-AI comparative cybersecurity research.
📝 Abstract
Human-agent evaluations often compress interaction into a single performance score, even when human and automated policies adapt differently over time. We study this in a sequential cyber-defense game on an attack graph, where a human or reinforcement-learning defender protects cloud assets against a Deep Q-Network (DQN) attacker. We compare four defender settings: a human reward-only game operationalizing the DQN information structure, a human reward-plus-transition game operationalizing the Cognitive Hierarchy Theory-driven DQN (CHT-DQN) information structure, an automated DQN defender, and an automated CHT-DQN defender. In the reward-only game, participants receive payoff and reward feedback. In the reward-plus-transition game, they also see attacker-aware transition probabilities from the CHT-DQN model. Across 80 Mechanical Turk participants and matched automated simulations, human defenders adapted to outcomes in a way the automated defenders did not: they were more likely to reselect a node after a successful defense than after a failure, in both games, an asymmetry we interpret as consistent with Prospect Theory and Cumulative Prospect Theory. The 40-round average also mixes early rounds, where the DQN attacker acts mostly at random, with late rounds, where it mostly exploits; in the final stage, mean protection ranks the reward-plus-transition human game above the reward-only human game, then the automated CHT-DQN defender, then the automated DQN defender, an ordering the overall mean hides. By contrast, the reward-plus-transition game does not produce a statistically reliable overall gain in weighted data protection over the reward-only game. These results suggest that evaluating human-agent cyber-defense systems only by an averaged task score can miss behaviorally meaningful differences in adaptation, action allocation, and bounded human reasoning.
Problem

Research questions and friction points this paper is trying to address.

Cyber Defense
Bounded Reasoning
Human-Agent Evaluation
Cognitive Hierarchy
Prospect Theory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cognitive Hierarchy Theory
Deep Q-Network
Cyber Defense
Bounded Reasoning
Prospect Theory