Score
Designs, generates, and evaluates adversarial textual suffixes—short appended strings intended to change a model's outputs—by creating candidate populations and measuring their effects via model queries. Builds and analyzes population-based optimization procedures (selection, mutation, crossover) and fitness-scoring mechanisms to discover and refine high‑fitness suffixes for a specified objective (e.g., inducing policy-violating responses).
This work investigates the robustness vulnerabilities of large language models (LLMs) against adversarial suffixes that bypass safety guard models—particularly Llama Prompt Guard 2—when processing untrusted inputs or generating executable code. We propose Super Suffixes, the first attack framework achieving cross-model and cross-tokenizer multi-objective alignment failure. Additionally, we design DeltaGuard, a lightweight detection mechanism that models malicious intent via cosine similarity to concept directions in residual streams. Experiments demonstrate that Super Suffixes successfully evade Prompt Guard 2’s protections across five mainstream generative models. DeltaGuard achieves a 99.8% malicious prompt detection rate, substantially enhancing guard model robustness. Crucially, this work is the first to expose Prompt Guard 2’s joint optimization vulnerability—where alignment and safety objectives are co-optimized—thereby establishing an interpretable, deployable paradigm for LLM safety.
Retrieval ranking in RAG systems is vulnerable to black-box adversarial prompt attacks, leading to erroneous generation. Method: This paper proposes a gradient-free differential evolution (DE) attack that generates highly stealthy, semantically natural adversarial suffixes—each ≤5 tokens—via a readability-aware suffix construction strategy, significantly reducing detection rates of MLM- and BERT-based detectors. Evaluation on the BEIR QA benchmark across diverse dense and sparse retrievers shows attack success rates comparable to or exceeding those of GGPP and PRADA, with near-random detection evasion rates. Contribution/Results: This work introduces DE into the black-box RAG attack framework for the first time and designs a lightweight suffix optimization mechanism that jointly balances stealthiness and efficacy. It establishes a novel paradigm for security evaluation of RAG systems.
Language models are vulnerable to short adversarial suffixes, yet existing gradient- or rule-based methods suffer from poor generalization and limited transferability across tasks and models. To address this, we propose the first reinforcement learning–based framework for generating universal adversarial suffixes: the suffix is modeled as a policy trained via proximal policy optimization (PPO); a calibrated cross-entropy reward mechanism mitigates label bias, while multi-task aggregation and reward shaping enhance cross-task and cross-model transferability; crucially, the target model’s parameters remain frozen throughout, with only sparse feedback derived from its output logits. Extensive experiments across five NLP benchmarks and three major language model families demonstrate that our method significantly degrades model accuracy, achieving higher attack success rates and superior transferability compared to state-of-the-art adversarial trigger techniques.
This work addresses the lack of systematic theoretical foundations for integrating large language models (LLMs) with evolutionary algorithms (EAs). At the microstructural level, it establishes, for the first time, a precise one-to-one mapping between their core components—identifying five fundamental analogies: token representation ↔ individual encoding, positional encoding ↔ fitness shaping, attention mechanisms ↔ selection pressure, layer-wise propagation ↔ generational transition, and parameter updates ↔ mutation operators. Methodologically, it introduces two novel paradigms: (i) *evolutionary fine-tuning*, where EAs guide LLM optimization via adaptive fitness design and cross-paradigm alignment; and (ii) *LLM-augmented EAs*, leveraging LLMs for population initialization, semantic-aware mutation, and surrogate fitness evaluation, grounded in Transformer architecture analysis. The contributions include uncovering latent evolutionary mechanisms inherent in LLMs and providing a scalable, interdisciplinary framework to enhance agent robustness, behavioral diversity, and continual learning capability.
Existing speculative decoding methods struggle to efficiently handle frequent, repetitive, and long-horizon predictable inference requests common in LLM-agent scenarios. This paper proposes SuffixDecoding—a novel, model-agnostic speculative decoding paradigm that operates entirely on CPU memory. Its core innovation is a dynamically maintained suffix tree over historical outputs, coupled with an interpretable, empirically calibrated token-frequency scoring mechanism to enable lightweight tree-based speculation and adaptive pruning. Compared to SpecInfer, SuffixDecoding achieves 1.4× higher throughput and 1.1× lower time-per-output-token (TPOT) latency in open-domain dialogue and code generation; in text-to-SQL tasks, it attains 2.9× throughput improvement and reduces latency to one-third, while sustaining high acceptance rates even under few-shot settings (256 examples). To our knowledge, this is the first speculative decoding framework that eliminates the need for a draft model, runs fully on CPU, and provides human-interpretable speculation decisions.
Despite the deployment of alignment and content moderation mechanisms, large language models (LLMs) remain vulnerable to black-box jailbreaking attacks. This work introduces, for the first time, a systematic application of genetic algorithms to black-box LLM jailbreaking, efficiently generating high-fitness adversarial suffixes by iteratively performing selection, mutation, and crossover operations in a discrete prompt space—without requiring access to internal model information. The proposed method successfully jailbreaks multiple mainstream commercial LLMs under realistic black-box conditions, substantially exposing the fragility of current safety safeguards and demonstrating the effectiveness and practical utility of evolution-inspired search strategies in adversarial prompt engineering.
Recently, Cenzato et al.\ proposed a new text index, called the \emph{suffixient array}, which is a subset of the suffix array and supports locating a single pattern occurrence or finding its maximal exact matches (MEMs), assuming random access to the input text $T[1..n]$ is available. They show that, given the suffix array, the longest common prefix array, and the Burrows--Wheeler transform (BWT) of the reverse of $T[1..n]$ over an alphabet $\{1,\ldots,σ\}$, a suffixient array can be constructed in linear time. However, their construction algorithms require multiple scans of these arrays. When restricted to a single pass over the arrays, they present an alternative construction algorithm running in $O(n + \overline{r} \log σ)$ time, where $\overline{r}$ is the number of runs in the BWT of the reversed text. In this paper, we present a new one-pass algorithm that constructs a suffixient array in linear time under the standard RAM model.
Natural language classifiers are vulnerable to semantics-preserving adversarial attacks in black-box settings, yet existing approaches suffer from limited efficiency and effectiveness. This work proposes GAversary, a novel method that, for the first time, integrates GloVe word embeddings into the mutation operator of a genetic algorithm to generate highly deceptive adversarial examples that maintain semantic similarity—requiring access only to the model’s logit outputs. Evaluated across multiple benchmark datasets, GAversary drastically reduces the target model’s accuracy from 76.8% to 5.8%, significantly outperforming state-of-the-art black-box attack methods such as BAE and A2T in terms of attack success rate.
This work addresses the challenge of effectively deciphering the evolutionary relationships, training lineages, and critical components of large language models (LLMs). It pioneers the systematic adaptation of phylogenetic inference methods from evolutionary biology to LLM analysis, drawing an analogy between model weights as genotypes and generated text as phenotypes. By constructing unsupervised phylogenetic trees, the approach reveals the lineage structure among models. Through weight-difference analysis, phenotypic experiments, and visualization, the method successfully reconstructs the true topology of training lineages, accurately identifies the network layers and training datasets that most significantly contribute to performance, and produces interpretable evolutionary relationship maps for multiple black-box foundation models.
Current safety alignment mechanisms in large language models are vulnerable to suffix-based GCG adversarial attacks and often overlook the influence of adversarial token placement. This work introduces token position as a critical variable into GCG attack analysis, systematically investigating its impact—particularly when adversarial tokens are placed in the prompt prefix—on attack success rates. By integrating position optimization with dynamic evaluation strategies, our experiments demonstrate that prefix-based attacks and strategic token repositioning significantly enhance adversarial effectiveness. These findings expose a critical blind spot in existing safety evaluation frameworks regarding positional sensitivity and advocate for a more comprehensive perspective on robustness assessment that explicitly accounts for token location within prompts.