🤖 AI Summary
This study addresses the high computational costs and model-pairing dependencies of existing adversarial attacks, which limit their scalability and transferability. To overcome these challenges, this work proposes ReSA, a novel framework that leverages maximum entropy inverse reinforcement learning to recover proxy reward models solely from behavioral observations, thereby exposing fundamental alignment vulnerabilities. Subsequently, it generates efficient adversarial strategies through reward-guided decoding inversion. Our experiments demonstrate that a single recovered reward function generalizes effectively across diverse models without requiring strict model pairing. Consequently, ReSA significantly outperforms existing methods in both attack effectiveness and transferability, offering a scalable approach to evaluating the robustness of aligned language models.
📝 Abstract
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement learning to recover a proxy reward model solely from the aligned model's behavior. The extracted reward is then reversed at inference time to derive an adversarial policy, efficiently implemented via a reward-guided decoding mechanism. Experiments demonstrate that a single recovered reward generalizes across prompts and diverse models to reveal a fundamental alignment vulnerability, enabling ReSA to significantly outperform existing attacks in effectiveness and transferability. The code is available at https://github.com/GarminQ/ReSA.