Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether reasoning models can conceal their computational processes under Chain-of-Thought (CoT) monitoring. Grounded in information-theoretic analysis and the Transformer architecture, this work introduces cryptographic assumptions to rigorously prove the separability between the existence of information-theoretic footprints and semantic readability. It reveals that while complex covert reasoning inevitably leaks informational traces, such traces can be rendered unreadable through encryption. By proposing a novel perspective on cryptographic evasion of monitoring, this research establishes fundamental theoretical boundaries for CoT oversight and provides a systematic analytical framework for AI safety evaluation.
📝 Abstract
Can reasoning models trick chain of thought (CoT) monitors and perform hidden computation without revealing it in their thinking traces? We show that the answer depends on the underlying task difficulty and the model size. Simple computations can be performed covertly; however, beyond a threshold depending on model size, successfully solving the task necessarily leaks a near-linear amount of information about the covert task input into the CoT. Therefore, sufficiently complex hidden computation always leaves an information-theoretic footprint. However, concerningly, this leakage need not be readable: Under plausible cryptographic assumptions, even a one-layer Transformer can encrypt its reasoning online so that no polynomial-time monitor can extract information about the hidden computation. Overall, our theoretical and empirical results provide a holistic view of both the opportunities and the limitations of CoT monitoring.
Problem

Research questions and friction points this paper is trying to address.

Chain-of-Thought monitoring
hidden reasoning
information leakage
AI safety
cryptographic obfuscation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chain-of-Thought Monitoring
Information Leakage
Hidden Reasoning
Cryptographic Assumptions
Transformer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mohammadali Mohammadkhani
Saarland University, Saarland Informatics Campus; Zuse School ELIZA
M
Madhava Krishna
Saarland University, Saarland Informatics Campus
Yash Sarrof
Yash Sarrof
Saarland University
Machine LearningLanguage
Michael Hahn
Michael Hahn
Saarland University
Machine LearningCognitionLanguage