Distillation Defenses Easily Break After Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of existing distillation defenses for large language models, which fail to protect against reasoning capability theft due to their neglect of subsequent reinforcement learning (RL) training. By revising the threat model, this work demonstrates that adversaries can replicate sophisticated attack outcomes using only simple API data combined with RL, rendering all current defense mechanisms that leak approximate reasoning trajectories ineffective. Building on this finding, the project proposes a novel batch-level defense paradigm, emphasizing the necessity of designing security mechanisms that do not expose internal reasoning details. This research fundamentally challenges conventional evaluation paradigms for defense effectiveness, providing critical theoretical foundations and new defensive directions for the intellectual property protection of large language models.
📝 Abstract
Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e.,"distill") their own models on these traces. Existing defenses against distillation attacks are typically evaluated immediately after distillation, implicitly assuming attackers do not train their models any further. In this paper, we argue that a more realistic threat model includes further training with reinforcement learning after distillation. A misspecified threat model can give a false sense of security -- some defenses that seem effective after distillation can be broken after subsequent reinforcement learning. Practically, reinforcement learning lowers the bar for a distillation attack to be effective. We show that simple attacks can steal reasoning capabilities from existing closed-source language models using data easily obtainable from current APIs, yielding reasoning improvements equivalent to more sophisticated attacks that extract the full hidden traces. Results indicate that any distillation defense that leaks sufficient information to reconstruct approximate reasoning traces is likely ineffective. We conclude by discussing broader implications and batch-level distillation defenses which could be more effective.
Problem

Research questions and friction points this paper is trying to address.

Distillation attack
Reinforcement learning
Defense robustness
Large language models
Threat model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distillation defense
Reinforcement learning
Threat model
Reasoning traces
Large language models
🔎 Similar Papers
No similar papers found.