🤖 AI Summary
This study addresses the overreliance of LLM agent evaluation on end-to-end success rates and the absence of multidimensional diagnostic criteria by proposing the concept of "cognitive depth," which encompasses five dimensions: contextual, temporal, multimodal, adaptive, and metacognitive. Methodologically, it constructs five-dimensional operationalized metrics alongside perturbation testing procedures to map agent capabilities onto world models. Furthermore, it incorporates techniques such as symbolic verifiers, structured memory, planning coupling, and tool constraints to enable trajectory-level diagnosis. This framework provides an item-wise diagnostic structure for benchmarks like GAIA, facilitating a fine-grained, quantitative assessment of agent capabilities.
📝 Abstract
Agentic large language model (LLM) systems are commonly implemented as an LLM in a loop with Planning, Memory, Tools, and Control Flow. This application-focused view connects agentic LLM research with deployable systems and leaves open how such systems should be evaluated beyond end-to-end task success. Building on this view, we define agentic cognitive depth as a trajectory-level profile across five operational criteria. The profile contains context sensitivity ($C$), temporal continuity ($T$), multimodal coordination ($M$), adaptive interaction ($A$), and metacognitive monitoring ($Mc$). The first four criteria measure how well Control Flow, Memory, Tools, and Planning are used across a trajectory. The fifth measures whether the system monitors and regulates the full run. For each criterion, we give operational proxies and a perturbation procedure, then connect the profile to the agent's world model. We provide the structure needed to extend benchmarks such as GAIA, SWE-bench, WebArena, and TRIP-Bench with per-criterion diagnostics. Symbolic verifiers, structured memory, planner coupling, and tool constraints provide practical ways to build and test these capacities.