🤖 AI Summary
This study addresses the vulnerability of large language model (LLM) agents to prompt injection and the limited generalizability of existing defenses against adaptive attacks by proposing SRFT, a Self-Reflective Fine-Tuning framework. Departing from static imitation paradigms, this work introduces a novel self-reflection mechanism grounded in failure experiences. Specifically, adversarial trajectories are constructed to expose agent vulnerabilities, while an expert model generates structured reflective reasoning to guide the agent in learning to identify malicious instructions and maintain alignment with user intent from past attack experiences. Experiments conducted on Llama-3.1 and Qwen3 demonstrate that SRFT significantly reduces attack success rates across both static and adaptive benchmarks while preserving task performance. The source code is publicly available.
📝 Abstract
Large language model (LLM) agents are increasingly deployed in tool-augmented environments, but their reliance on external inputs makes them highly vulnerable to prompt injection attacks that can hijack task objectives. Existing safety alignment methods rely on static expert trajectories or preference optimization, limiting their ability to generalize to adaptive attack patterns. In this work, we propose Self-Reflection Fine-Tuning (SRFT), a training framework that enables agents to improve robustness by learning from their own failure experiences under adversarial conditions. Instead of passively imitating expert behaviors, SRFT exposes the agent to compromised trajectories constructed via injected attacks, and leverages an expert model to generate structured self-reflection reasoning that contrasts unsafe and optimal actions. This reflective supervision teaches the agent to identify malicious instructions, reason about their consequences, and maintain alignment with the original user intent. We instantiate this framework in SR-Agent, built on Llama-3.1-8B-Instruct and Qwen3-8B, and evaluate it on both static and adaptive prompt injection benchmarks. Experimental results show that SRFT substantially reduces attack success rates while preserving task performance, and demonstrates strong generalization under adaptive attacks. These findings suggest that learning from failure via self-reflection is a promising direction for building robust and secure LLM agents. Our code is released at https://github.com/Eden-Wang1710/srft-repo.