🤖 AI Summary
This study addresses the significant safety degradation in existing token-level adaptive computation models caused by layer-skipping mechanisms. We propose SafeDepth, a lightweight plug-in framework that first reveals how layer skipping compromises safety and distinguishes the computational depth required for harmful input refusal versus benign query identification. Specifically, SafeDepth employs a dynamic router to select computation paths based on token, layer position, and processing stage, complemented by adapter modules that compensate for information lost from skipped layers. The router and adapters are jointly trained with a frozen backbone to optimize the safety-efficiency trade-off. Experiments on Llama-3-8B demonstrate that SafeDepth effectively reduces computational overhead while decreasing the harmful response rate on HarmBench-HJ from 55.44% to 25.25% and the over-refusal rate on XSTest from 12.0% to 4.0%, all without compromising general task performance.
📝 Abstract
Recent studies suggest that not every token needs to pass through all Transformer layers, motivating token-level adaptive models that selectively skip layers to reduce computation. Our experiments show that these execution choices also affect safety: existing token-level adaptive reasoning models exhibit higher harmful-response rates than their reference backbones. The safety effects depend on which layers are skipped and whether skipping occurs during prompt processing or answer generation. We further find that recognizing harmful requests and producing refusals depend on different layer-wise computations. We introduce SafeDepth, a lightweight, plug-in framework that uses selective layer execution to improve both safety and efficiency. SafeDepth learns to retain computations that support safe responses and bypass those that contribute to unsafe generation. A router selects layer execution based on token representations, layer position, and inference phase, while an adapter supports the skipped paths. We jointly train these modules to reduce computation and harmful responses while preserving performance on benign tasks, targeting a better safety-efficiency trade-off. The pretrained backbone remains frozen throughout, requiring no additional pretraining. Experimental results on Llama-3-8B-Instruct show that SafeDepth reduces computation relative to full-depth inference while largely preserving task performance. Compared with FlexiDepth, it lowers unsafe-response rates across all five harmful-request benchmarks, with the largest absolute decrease on HarmBench-HJ (from 55.44% to 25.25%), while reducing the XSTest false-refusal rate from 12.0% to 4.0%.