adversarial suffix evolution

Designs, generates, and evaluates adversarial textual suffixes—short appended strings intended to change a model's outputs—by creating candidate populations and measuring their effects via model queries. Builds and analyzes population-based optimization procedures (selection, mutation, crossover) and fitness-scoring mechanisms to discover and refine high‑fitness suffixes for a specified objective (e.g., inducing policy-violating responses).

adversarialsuffixevolution

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously

Dec 12, 2025
AA
Andrew Adiletta
🏛️ MITRE | Worcester Polytechnic Institute

This work investigates the robustness vulnerabilities of large language models (LLMs) against adversarial suffixes that bypass safety guard models—particularly Llama Prompt Guard 2—when processing untrusted inputs or generating executable code. We propose Super Suffixes, the first attack framework achieving cross-model and cross-tokenizer multi-objective alignment failure. Additionally, we design DeltaGuard, a lightweight detection mechanism that models malicious intent via cosine similarity to concept directions in residual streams. Experiments demonstrate that Super Suffixes successfully evade Prompt Guard 2’s protections across five mainstream generative models. DeltaGuard achieves a 99.8% malicious prompt detection rate, substantially enhancing guard model robustness. Crucially, this work is the first to expose Prompt Guard 2’s joint optimization vulnerability—where alignment and safety objectives are co-optimized—thereby establishing an interpretable, deployable paradigm for LLM safety.

Bypassing guard models for malicious text generationCompromising Llama Prompt Guard 2 via joint optimizationDetecting Super Suffix attacks using model internal states

Retrieval ranking in RAG systems is vulnerable to black-box adversarial prompt attacks, leading to erroneous generation. Method: This paper proposes a gradient-free differential evolution (DE) attack that generates highly stealthy, semantically natural adversarial suffixes—each ≤5 tokens—via a readability-aware suffix construction strategy, significantly reducing detection rates of MLM- and BERT-based detectors. Evaluation on the BEIR QA benchmark across diverse dense and sparse retrievers shows attack success rates comparable to or exceeding those of GGPP and PRADA, with near-random detection evasion rates. Contribution/Results: This work introduces DE into the black-box RAG attack framework for the first time and designs a lightweight suffix optimization mechanism that jointly balances stealthiness and efficacy. It establishes a novel paradigm for security evaluation of RAG systems.

Attacking RAG systems via adversarial prompt injectionEvading detection while maintaining readabilityOptimizing prompts to manipulate retrieval rankings

Universal Adversarial Suffixes for Language Models Using Reinforcement Learning with Calibrated Reward

Dec 08, 2025
SS
Sampriti Soor
🏛️ Indian Institute of Technology Guwahati

Language models are vulnerable to short adversarial suffixes, yet existing gradient- or rule-based methods suffer from poor generalization and limited transferability across tasks and models. To address this, we propose the first reinforcement learning–based framework for generating universal adversarial suffixes: the suffix is modeled as a policy trained via proximal policy optimization (PPO); a calibrated cross-entropy reward mechanism mitigates label bias, while multi-task aggregation and reward shaping enhance cross-task and cross-model transferability; crucially, the target model’s parameters remain frozen throughout, with only sparse feedback derived from its output logits. Extensive experiments across five NLP benchmarks and three major language model families demonstrate that our method significantly degrades model accuracy, achieving higher attack success rates and superior transferability compared to state-of-the-art adversarial trigger techniques.

Develops universal adversarial suffixes using reinforcement learningEvaluates method on diverse NLP benchmarks with multiple language modelsImproves transferability across tasks and models via calibrated rewards

This work addresses the lack of systematic theoretical foundations for integrating large language models (LLMs) with evolutionary algorithms (EAs). At the microstructural level, it establishes, for the first time, a precise one-to-one mapping between their core components—identifying five fundamental analogies: token representation ↔ individual encoding, positional encoding ↔ fitness shaping, attention mechanisms ↔ selection pressure, layer-wise propagation ↔ generational transition, and parameter updates ↔ mutation operators. Methodologically, it introduces two novel paradigms: (i) *evolutionary fine-tuning*, where EAs guide LLM optimization via adaptive fitness design and cross-paradigm alignment; and (ii) *LLM-augmented EAs*, leveraging LLMs for population initialization, semantic-aware mutation, and surrogate fitness evaluation, grounded in Transformer architecture analysis. The contributions include uncovering latent evolutionary mechanisms inherent in LLMs and providing a scalable, interdisciplinary framework to enhance agent robustness, behavioral diversity, and continual learning capability.

Analyzing evolutionary fine-tuning and LLM-enhanced EAs challenges.Exploring parallels between LLMs and EAs for mutual enhancement.Providing insights into evolutionary mechanisms behind LLMs.

SuffixDecoding: A Model-Free Approach to Speeding Up Large Language Model Inference

Nov 07, 2024
GO
Gabriele Oliaro
🏛️ Carnegie Mellon University | Snowflake AI Research

Existing speculative decoding methods struggle to efficiently handle frequent, repetitive, and long-horizon predictable inference requests common in LLM-agent scenarios. This paper proposes SuffixDecoding—a novel, model-agnostic speculative decoding paradigm that operates entirely on CPU memory. Its core innovation is a dynamically maintained suffix tree over historical outputs, coupled with an interpretable, empirically calibrated token-frequency scoring mechanism to enable lightweight tree-based speculation and adaptive pruning. Compared to SpecInfer, SuffixDecoding achieves 1.4× higher throughput and 1.1× lower time-per-output-token (TPOT) latency in open-domain dialogue and code generation; in text-to-SQL tasks, it attains 2.9× throughput improvement and reduces latency to one-third, while sustaining high acceptance rates even under few-shot settings (256 examples). To our knowledge, this is the first speculative decoding framework that eliminates the need for a draft model, runs fully on CPU, and provides human-interpretable speculation decisions.

Enhancing LLM inference speed via adaptive suffix tree cachingExploiting repetitive inference requests in agentic AI workloadsImproving speculative decoding for long predictable sequences

Latest Papers

What's happening recently
View more

Despite the deployment of alignment and content moderation mechanisms, large language models (LLMs) remain vulnerable to black-box jailbreaking attacks. This work introduces, for the first time, a systematic application of genetic algorithms to black-box LLM jailbreaking, efficiently generating high-fitness adversarial suffixes by iteratively performing selection, mutation, and crossover operations in a discrete prompt space—without requiring access to internal model information. The proposed method successfully jailbreaks multiple mainstream commercial LLMs under realistic black-box conditions, substantially exposing the fragility of current safety safeguards and demonstrating the effectiveness and practical utility of evolution-inspired search strategies in adversarial prompt engineering.

adversarial attackblack-box settingLLM jailbreaking

Recently, Cenzato et al.\ proposed a new text index, called the \emph{suffixient array}, which is a subset of the suffix array and supports locating a single pattern occurrence or finding its maximal exact matches (MEMs), assuming random access to the input text $T[1..n]$ is available. They show that, given the suffix array, the longest common prefix array, and the Burrows--Wheeler transform (BWT) of the reverse of $T[1..n]$ over an alphabet $\{1,\ldots,σ\}$, a suffixient array can be constructed in linear time. However, their construction algorithms require multiple scans of these arrays. When restricted to a single pass over the arrays, they present an alternative construction algorithm running in $O(n + \overline{r} \log σ)$ time, where $\overline{r}$ is the number of runs in the BWT of the reversed text. In this paper, we present a new one-pass algorithm that constructs a suffixient array in linear time under the standard RAM model.

Burrows-Wheeler transformlinear timesingle-pass construction

Natural language classifiers are vulnerable to semantics-preserving adversarial attacks in black-box settings, yet existing approaches suffer from limited efficiency and effectiveness. This work proposes GAversary, a novel method that, for the first time, integrates GloVe word embeddings into the mutation operator of a genetic algorithm to generate highly deceptive adversarial examples that maintain semantic similarity—requiring access only to the model’s logit outputs. Evaluated across multiple benchmark datasets, GAversary drastically reduces the target model’s accuracy from 76.8% to 5.8%, significantly outperforming state-of-the-art black-box attack methods such as BAE and A2T in terms of attack success rate.

adversarial textblack-box attacksnatural language classifiers

This work addresses the challenge of effectively deciphering the evolutionary relationships, training lineages, and critical components of large language models (LLMs). It pioneers the systematic adaptation of phylogenetic inference methods from evolutionary biology to LLM analysis, drawing an analogy between model weights as genotypes and generated text as phenotypes. By constructing unsupervised phylogenetic trees, the approach reveals the lineage structure among models. Through weight-difference analysis, phenotypic experiments, and visualization, the method successfully reconstructs the true topology of training lineages, accurately identifies the network layers and training datasets that most significantly contribute to performance, and produces interpretable evolutionary relationship maps for multiple black-box foundation models.

evolutionary analysisexplainabilitylarge language models

Current safety alignment mechanisms in large language models are vulnerable to suffix-based GCG adversarial attacks and often overlook the influence of adversarial token placement. This work introduces token position as a critical variable into GCG attack analysis, systematically investigating its impact—particularly when adversarial tokens are placed in the prompt prefix—on attack success rates. By integrating position optimization with dynamic evaluation strategies, our experiments demonstrate that prefix-based attacks and strategic token repositioning significantly enhance adversarial effectiveness. These findings expose a critical blind spot in existing safety evaluation frameworks regarding positional sensitivity and advocate for a more comprehensive perspective on robustness assessment that explicitly accounts for token location within prompts.

adversarial attacksGCGjailbreak attacks

Hot Scholars

SI

Shunsuke Inenaga

Professor, Department of Informatics, Kyushu University
Algorithms and Data StructuresString AlgorithmsCompressionCombinatorics on Words
TM

Takuya Mieno

The University of Electro-Communications
Stringology
WD

Wenbin Dai

Shanghai Jiao Tong University
Industrial Edge ComputingIndustrial InformaticsAutomation Code GenerationIndustrial Control Software
LF

Luigi Foscari

University of Milan
online learninggame theorymulti-agent reinforcement learning