Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks

📅 2026-06-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of traditional red-teaming evaluations that rely solely on attack success rate (ASR) by introducing process mining to analyze the temporal dynamics of large language models during adversarial interactions. Leveraging 8,575 annotated events, the authors construct direct-follow graphs and state transition matrices to uncover dynamic defense mechanisms. Their analysis reveals that GPT-OSS exhibits strong refusal behavior akin to an absorbing state, whereas Llama models display multiple vulnerable pathways susceptible to exploitation. Furthermore, significant differences emerge across models in terms of mutator efficacy and jailbreak time distributions. By moving beyond static ASR metrics, this approach elucidates structural disparities in defensive strategies among models, offering a more nuanced understanding of their robustness against adversarial attacks.
📝 Abstract
Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate (ASR), not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process models from event logs, to red teaming traces. We conduct a controlled experiment pitting 60 HarmBench prompts against two LLMs, GPT-OSS 120B and Llama 3.3 70B, using 10 prompt mutation strategies over up to 110 attempts per prompt. From the resulting 8,575 scored events we extract Directly-Follows Graphs (DFGs) and state transition matrices that reveal structurally distinct defense profiles invisible to ASR alone: GPT-OSS exhibits a near-absorbing refusal state, while Llama presents multiple porous escape routes from refusal to getting successfully jailbroken. We further show that mutator effectiveness is asymmetric across models and that time-to-jailbreak distributions differ by an order of magnitude.
Problem

Research questions and friction points this paper is trying to address.

red teaming
attack success rate
process mining
LLM security
adversarial robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

process mining
red teaming
Directly-Follows Graph
state transition matrix
jailbreak resistance