Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether model behavior can be effectively monitored in latent chain-of-thought (CoT) reasoning—where explicit, human-readable reasoning traces are absent. Through prompt interventions, activation probing, and latent state textualization, the authors systematically evaluate monitoring efficacy across mathematical reasoning and question-answering tasks under explicit CoT and both weakly and strongly supervised latent CoT settings. The work reveals, for the first time, that monitoring performance depends primarily on the degree of constraint imposed by the task’s correct answers and the extent of internal model access, rather than on the presence or absence of explicit reasoning chains. Notably, effective monitoring remains achievable even without explicit CoT, provided sufficient access to the model’s internal representations is available.
📝 Abstract
Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors by reading the reasoning trace, motivating work on CoT monitorability. Latent CoT approaches, however, replace the explicit tokens with a small number of continuous states, lowering inference costs but removing the readable trace this monitoring relies on. Monitoring then requires alternative access to the model, such as probing its activations or verbalizing the latent states back into text, but how much monitorability these alternatives preserve is unclear. We study this question with a hint-based intervention setup, a proxy for behaviors where models exploit biasing input cues, e.g., an inadvertently leaked answer or a belief stated by the user, without acknowledging them. Taking hint-reliance as the monitorability target, we compare monitors across reasoning modes, from explicit CoT to weakly- and strongly-supervised latent CoT, on math reasoning and question answering. We find that, in this setup, monitorability depends more on properties of the task (such as whether the correct answer constrains the supporting reasoning) and the level of access to model internals than on the reasoning mode.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought
monitorability
latent reasoning
large language models
hint-based intervention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought
Monitorability
Latent CoT
Hint-based Intervention
Model Interpretability
🔎 Similar Papers
No similar papers found.