🤖 AI Summary
This work proposes a “layered attribution” diagnostic framework to disentangle the origins of inscrutable behaviors exhibited by AI agents in complex social systems, which are often conflated between internal representations and external constraints. The framework systematically distinguishes a foundational computational layer—encompassing architecture, memory, and perception—from a behavioral modulation layer comprising identity, goals, social interactions, and institutional constraints, thereby integrating representation learning, multi-agent modeling, and institutional analysis into a unified two-tier diagnostic architecture. It yields three key insights: behavioral substitutability validity hinges on the coupling among model, task, and layer; human–AI behavioral discrepancies can serve as diagnostic signals; and effective governance presupposes precise source attribution. This approach establishes a theoretical foundation for interpreting and governing AI behavior.
📝 Abstract
AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating these systems requires source-level diagnosis: the same behavioral pattern may arise from an agent representational substrate or from the roles, objectives, interaction structures, and governance rules that shape its expression. This Perspective proposes a diagnostic framework for AI agent behavior: layer attribution. The foundational computational layer defines what behaviors are possible through architecture, memory, perception, attention, and representation. The behavioral modulation layer shapes how those capacities are expressed through identity, resources, objectives, social interaction, institutional constraints, and governance. The framework clarifies three consequences: surrogate validity is a model-task-layer relation, human-AI divergence provides diagnostic evidence, and governance requires source attribution before intervention. Treating AI agents as behavioral actors therefore requires evaluation methods that determine where behavior originates before deciding how to explain, validate, or govern it.