Institution profile

SAIGE

Research institution
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard

Sep 30, 2026

This study addresses the challenge that steganographic reasoning in large language models is difficult to supervise externally and its acquisition mechanisms remain unclear. We systematically compare three paradigms—reinforcement learning, in-context learning, and supervised fine-tuning (SFT)—in enabling models to acquire steganographic reasoning, while investigating its intrinsic relationship with steganographic communication and coded reasoning. Our findings reveal that the learning threshold for steganographic reasoning is significantly higher than that of basic steganography, and mastering the latter does not necessarily lead to the emergence of the former. This capability is typically acquired only through SFT with slow convergence, yet can be rapidly mastered when cover tasks are facilitative. By exposing the high learning barrier of steganographic reasoning, this work provides critical empirical evidence for the safety evaluation of covert behaviors in large language models.

0 citationsRead paper
Recent publications

Latest Papers

Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard

Sep 30, 2026

This study addresses the challenge that steganographic reasoning in large language models is difficult to supervise externally and its acquisition mechanisms remain unclear. We systematically compare three paradigms—reinforcement learning, in-context learning, and supervised fine-tuning (SFT)—in enabling models to acquire steganographic reasoning, while investigating its intrinsic relationship with steganographic communication and coded reasoning. Our findings reveal that the learning threshold for steganographic reasoning is significantly higher than that of basic steganography, and mastering the latter does not necessarily lead to the emergence of the former. This capability is typically acquired only through SFT with slow convergence, yet can be rapidly mastered when cover tasks are facilitative. By exposing the high learning barrier of steganographic reasoning, this work provides critical empirical evidence for the safety evaluation of covert behaviors in large language models.

0 citationsRead paper