🤖 AI Summary
This work addresses the vulnerability of large language models acting as autonomous agents in complex tool-augmented environments, where contextual contamination and over-reasoning often arise due to the absence of effective mechanisms for evaluating both their own capabilities and the reliability of external tools. To tackle this, the paper introduces the MESA-S framework, which, for the first time in a single-agent setting, decouples skill utility awareness from execution. By integrating metacognitive skill cards, delayed process probes, and a dual-dimensional confidence model, MESA-S computationally instantiates human-like cognitive control mechanisms—including delayed evaluation, cognitive vigilance, and proximal offloading—and formalizes trust provenance. Experimental results demonstrate that this approach effectively mitigates supply-chain vulnerabilities, curbs confidence inflation, reduces redundant reasoning, and significantly enhances agent behavioral reliability.
📝 Abstract
As large language models (LLMs) transition into autonomous agents integrated with extensive tool ecosystems, traditional routing heuristics increasingly succumb to context pollution and "overthinking". We argue that the bottleneck is not a deficit in algorithmic capability or skill diversity, but the absence of disciplined second-order metacognitive governance. In this paper, our scientific contribution focuses on the computational translation of human cognitive control - specifically, delayed appraisal, epistemic vigilance, and region-of-proximal offloading - into a single-agent architecture. We introduce MESA-S (Metacognitive Skills for Agents, Single-agent), a preliminary framework that shifts scalar confidence estimation into a vector separating self-confidence (parametric certainty) from source-confidence (trust in retrieved external procedures). By formalizing a delayed procedural probe mechanism and introducing Metacognitive Skill Cards, MESA-S decouples the awareness of a skill's utility from its token-intensive execution. Evaluated under an In-Context Static Benchmark Evaluation natively executed via Gemini 3.1 Pro, our early results suggest that explicitly programming trust provenance and delayed escalation mitigates supply-chain vulnerabilities, prunes unnecessary reasoning loops, and prevents offloading-induced confidence inflation. This architecture offers a scientifically cautious, behaviorally anchored step toward reliable, epistemically vigilant single-agent orchestration.