From Expected Harmfulness to Likelihood: A Probabilistic Reformulation of Jailbreaking LLM Agents

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the long-standing theoretical disconnect in LLM jailbreak research between two optimization perspectives: maximizing expected harm and increasing target output likelihood. Through a probabilistic reformulation, this work reveals their intrinsic unification by proving gradient equivalence, and accordingly proposes the OPUR sampling distribution, which leverages reweighted output distributions to generate high-harm target sequences and guide gradient-based input optimization. This research provides the first theoretical reconciliation of these divergent jailbreaking paradigms. The proposed OPUR method enables efficiently guided jailbreak attacks, with experiments demonstrating its significant effectiveness in circumventing the safety constraints of LLM agents.
📝 Abstract
When the harmfulness of an LLM agent's output can be quantified, a natural jailbreaking objective is to maximize expected harmfulness over admissible input modifications. An alternative approach constructs or selects harmful target outputs and modifies the input to increase their likelihood. We establish a precise connection between these two approaches through a probabilistic reformulation. Specifically, we show that the gradient of the logarithm of expected harmfulness with respect to the input equals the expected input gradient of the model's log-likelihood under a harmfulness reweighted output distribution. This identity provides a unified interpretation of expected harmfulness and target likelihood optimization. Building on this connection, we propose OPUR, a sampling distribution designed to generate highly harmful target outputs and use the resulting samples to guide likelihood-based input optimization. Experiments demonstrate the effectiveness of the resulting method in jailbreaking LLM agents.
Problem

Research questions and friction points this paper is trying to address.

jailbreaking
LLM agents
expected harmfulness
target likelihood
probabilistic reformulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Jailbreaking LLM Agents
Probabilistic Reformulation
Expected Harmfulness
Likelihood Optimization
OPUR
🔎 Similar Papers
No similar papers found.