Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the oversight of cognitive humility in existing LLM agent evaluations under knowledge conflicts. We propose a trajectory-level cognitive humility evaluation framework that systematically examines agents' uncertainty management when parametric knowledge conflicts with external evidence across three dimensions: recognition, resolution, and escalation. Through controlled and naturally occurring knowledge conflict tests coupled with multi-step execution trajectory analysis, we reveal a divergence between high model accuracy and low cognitive humility. Our findings demonstrate that while model interventions can enhance humility, they do so at the cost of accuracy, confirming that both properties emerge from systemic interactions. This work offers a novel perspective for developing more reliable LLM agents by highlighting the inherent trade-offs between performance and epistemic calibration.
📝 Abstract
When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.
Problem

Research questions and friction points this paper is trying to address.

Epistemic Humility
Knowledge Conflict
LLM Agents
Uncertainty Communication
Innovation

Methods, ideas, or system contributions that make the work stand out.

Epistemic Humility
Knowledge Conflict
LLM Agents
Trajectory-level Evaluation
ISE Framework
🔎 Similar Papers
No similar papers found.