SafeDepth: Safety-Aware Token-Level Adaptive Computation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant safety degradation in existing token-level adaptive computation models caused by layer-skipping mechanisms. We propose SafeDepth, a lightweight plug-in framework that first reveals how layer skipping compromises safety and distinguishes the computational depth required for harmful input refusal versus benign query identification. Specifically, SafeDepth employs a dynamic router to select computation paths based on token, layer position, and processing stage, complemented by adapter modules that compensate for information lost from skipped layers. The router and adapters are jointly trained with a frozen backbone to optimize the safety-efficiency trade-off. Experiments on Llama-3-8B demonstrate that SafeDepth effectively reduces computational overhead while decreasing the harmful response rate on HarmBench-HJ from 55.44% to 25.25% and the over-refusal rate on XSTest from 12.0% to 4.0%, all without compromising general task performance.
📝 Abstract
Recent studies suggest that not every token needs to pass through all Transformer layers, motivating token-level adaptive models that selectively skip layers to reduce computation. Our experiments show that these execution choices also affect safety: existing token-level adaptive reasoning models exhibit higher harmful-response rates than their reference backbones. The safety effects depend on which layers are skipped and whether skipping occurs during prompt processing or answer generation. We further find that recognizing harmful requests and producing refusals depend on different layer-wise computations. We introduce SafeDepth, a lightweight, plug-in framework that uses selective layer execution to improve both safety and efficiency. SafeDepth learns to retain computations that support safe responses and bypass those that contribute to unsafe generation. A router selects layer execution based on token representations, layer position, and inference phase, while an adapter supports the skipped paths. We jointly train these modules to reduce computation and harmful responses while preserving performance on benign tasks, targeting a better safety-efficiency trade-off. The pretrained backbone remains frozen throughout, requiring no additional pretraining. Experimental results on Llama-3-8B-Instruct show that SafeDepth reduces computation relative to full-depth inference while largely preserving task performance. Compared with FlexiDepth, it lowers unsafe-response rates across all five harmful-request benchmarks, with the largest absolute decrease on HarmBench-HJ (from 55.44% to 25.25%), while reducing the XSTest false-refusal rate from 12.0% to 4.0%.
Problem

Research questions and friction points this paper is trying to address.

token-level adaptive computation
safety
selective layer skipping
harmful responses
efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-level Adaptive Computation
Safety-Aware Routing
Selective Layer Execution
Plug-in Framework
Safety-Efficiency Trade-off
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
N
Nizhang Li
Faculty of Innovation Engineering, Macau University of Science and Technology
Ian G. Harris
Ian G. Harris
University of California Irvine
Design VerificationComputer SecurityNatural Language Processing