Reward Stealing Attack on Large Language Models

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational costs and model-pairing dependencies of existing adversarial attacks, which limit their scalability and transferability. To overcome these challenges, this work proposes ReSA, a novel framework that leverages maximum entropy inverse reinforcement learning to recover proxy reward models solely from behavioral observations, thereby exposing fundamental alignment vulnerabilities. Subsequently, it generates efficient adversarial strategies through reward-guided decoding inversion. Our experiments demonstrate that a single recovered reward function generalizes effectively across diverse models without requiring strict model pairing. Consequently, ReSA significantly outperforms existing methods in both attack effectiveness and transferability, offering a scalable approach to evaluating the robustness of aligned language models.
📝 Abstract
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.
Problem

Research questions and friction points this paper is trying to address.

Adversarial Attack
Large Language Models
Transferability
Scalability
Safety Alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward Stealing Attack
Inverse Reinforcement Learning
Adversarial Attack
LLM Alignment
Transferability
🔎 Similar Papers