Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
This study addresses the challenge that steganographic reasoning in large language models is difficult to supervise externally and its acquisition mechanisms remain unclear. We systematically compare three paradigms—reinforcement learning, in-context learning, and supervised fine-tuning (SFT)—in enabling models to acquire steganographic reasoning, while investigating its intrinsic relationship with steganographic communication and coded reasoning. Our findings reveal that the learning threshold for steganographic reasoning is significantly higher than that of basic steganography, and mastering the latter does not necessarily lead to the emergence of the former. This capability is typically acquired only through SFT with slow convergence, yet can be rapidly mastered when cover tasks are facilitative. By exposing the high learning barrier of steganographic reasoning, this work provides critical empirical evidence for the safety evaluation of covert behaviors in large language models.