Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents

📅 2026-05-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether the behavior of large language model (LLM) agents aligns with their stated reasoning processes—a property termed “process fidelity.” To this end, we introduce a novel experimental framework within the controlled social simulation of Texas Hold’em poker that requires no reference to ground-truth behaviors. We decompose the “faithfulness gap” into two quantifiable stages: reasoning-to-conclusion and conclusion-to-action. Our analysis reveals opposing trends across these stages, exposing a significant disconnect between LLM agents’ explicit reasoning and their actual decisions. These findings offer a new perspective and methodological foundation for evaluating the internal consistency of LLM-based agents in strategic, multi-agent settings.
📝 Abstract
Do LLM agents act on the reasoning they state? This question of process fidelity is central to using LLMs in social simulation, yet it is hard to measure where no reference for correct behavior exists. We study it in acontrolled setting, a Texas Poker simulator with a verifiable reference action for every decision by decomposing the faithfulness gap into two steps: reasoning-conclusion and conclusion-action. The two steps behave oppositely.
Problem

Research questions and friction points this paper is trying to address.

faithfulness gap
LLM agents
process fidelity
reasoning-action alignment
social simulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

faithfulness gap
process fidelity
LLM agents
reasoning-action alignment
controlled simulation