🤖 AI Summary
This study addresses the lack of a unified explanatory framework for jailbreak attacks on diffusion-based large models by formulating safety alignment as the shaping of a denoising energy landscape. It proposes complementary detection signals based on initial safety and trajectory kinetic energy. By constructing an energy landscape analysis framework, this work reveals the fundamental mechanisms underlying such attacks, demonstrating that adversaries must either overcome energy barriers or expose malicious intent. Accordingly, detection is achieved by integrating log-distribution analysis with trajectory velocity monitoring. Experimental results across multiple models validate the complementarity of these signals, confirming that the proposed approach enables efficient, training-free detection that remains robust against evasion attempts.
📝 Abstract
Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary, training-free detection signals: a step-0 ratio that reads the initial safety disposition from the logit distribution before generation begins, and two trajectory-velocity signals that track kinetic energy in complementary subspaces of the logit space. An attack must either reveal its intent at initialisation or expend kinetic energy to cross the barrier in at least one monitored subspace, so the three signals cover each other's blind spots in the energy budget by construction. Evaluation across three dense dLLMs (LLaDA-8B, LLaDA-1.5, Dream-7B) and a sparse mixture-of-experts dLLM (LLaDA-MoE-7B) confirms this complementarity. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that the detection and barrier-crossing thresholds are hard to separate.