🤖 AI Summary
This study investigates whether the behavior of large language model (LLM) agents aligns with their stated reasoning processes—a property termed “process fidelity.” To this end, we introduce a novel experimental framework within the controlled social simulation of Texas Hold’em poker that requires no reference to ground-truth behaviors. We decompose the “faithfulness gap” into two quantifiable stages: reasoning-to-conclusion and conclusion-to-action. Our analysis reveals opposing trends across these stages, exposing a significant disconnect between LLM agents’ explicit reasoning and their actual decisions. These findings offer a new perspective and methodological foundation for evaluating the internal consistency of LLM-based agents in strategic, multi-agent settings.
📝 Abstract
Do LLM agents act on the reasoning they state? This question of process fidelity is central to using LLMs in social simulation, yet it is hard to measure where no reference for correct behavior exists. We study it in acontrolled setting, a Texas Poker simulator with a verifiable reference action for every decision by decomposing the faithfulness gap into two steps: reasoning-conclusion and conclusion-action. The two steps behave oppositely.