Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs

📅 2026-04-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of large language models acting as autonomous agents in complex tool-augmented environments, where contextual contamination and over-reasoning often arise due to the absence of effective mechanisms for evaluating both their own capabilities and the reliability of external tools. To tackle this, the paper introduces the MESA-S framework, which, for the first time in a single-agent setting, decouples skill utility awareness from execution. By integrating metacognitive skill cards, delayed process probes, and a dual-dimensional confidence model, MESA-S computationally instantiates human-like cognitive control mechanisms—including delayed evaluation, cognitive vigilance, and proximal offloading—and formalizes trust provenance. Experimental results demonstrate that this approach effectively mitigates supply-chain vulnerabilities, curbs confidence inflation, reduces redundant reasoning, and significantly enhances agent behavioral reliability.

Technology Category

Multiagent Systems: Adversarial AgentsCognitive Modeling & Cognitive Systems: Agent ArchitecturesNatural Language Processing: Safety and Robustness

Application Category

Responsible Web: Machine-in-the-loop, human agency and autonomySearch and Retrieval-Augmented AI: Agentic searchSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
As large language models (LLMs) transition into autonomous agents integrated with extensive tool ecosystems, traditional routing heuristics increasingly succumb to context pollution and "overthinking". We argue that the bottleneck is not a deficit in algorithmic capability or skill diversity, but the absence of disciplined second-order metacognitive governance. In this paper, our scientific contribution focuses on the computational translation of human cognitive control - specifically, delayed appraisal, epistemic vigilance, and region-of-proximal offloading - into a single-agent architecture. We introduce MESA-S (Metacognitive Skills for Agents, Single-agent), a preliminary framework that shifts scalar confidence estimation into a vector separating self-confidence (parametric certainty) from source-confidence (trust in retrieved external procedures). By formalizing a delayed procedural probe mechanism and introducing Metacognitive Skill Cards, MESA-S decouples the awareness of a skill's utility from its token-intensive execution. Evaluated under an In-Context Static Benchmark Evaluation natively executed via Gemini 3.1 Pro, our early results suggest that explicitly programming trust provenance and delayed escalation mitigates supply-chain vulnerabilities, prunes unnecessary reasoning loops, and prevents offloading-induced confidence inflation. This architecture offers a scientifically cautious, behaviorally anchored step toward reliable, epistemically vigilant single-agent orchestration.
Problem

Research questions and friction points this paper is trying to address.

epistemic vigilance
delayed appraisal
trust calibration
single-agent LLMs
metacognitive governance
Innovation

Methods, ideas, or system contributions that make the work stand out.

delayed appraisal
epistemic vigilance
metacognitive governance
confidence decomposition
single-agent LLMs
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Eren Unlu
Globeholder, Paris, France