Score
Design, implement, and analyze evaluation procedures that measure a model's ability to reproduce exact token sequences (exact recall), including greedy decoding tests, behavioral secret-extraction experiments, and hit@k-style metrics that check whether a target k-token sequence appears among a model's top outputs; also develop protocols to compute and report hit@k and to separate behavioral output-based recall from probabilistic probes such as NLL.
This study addresses a critical flaw in existing security evaluations, wherein stored labels are erroneously treated as ground-truth behavioral facts, leading to processing leaks and construct validity threats. To rectify this, the authors propose a seven-stage integrity chain and an endpoint integrity checker that jointly audit historical label assignment mechanisms through execution trace tracing, blind processing reconstruction, double-blind consistency review, and semantic boundary analysis. This approach exposes the non-terminal nature of labels and reconstructs the measurement foundation of security assessments. Empirically, the method corrects 58 mislabeled instances, eliminates all false-positive attack records in the v2 census while preserving genuine privilege escalation cases, and achieves consensus across reviewers on 96 explainable requests.
Traditional language model evaluation often conflates the ability to produce assessable responses with the correctness of those responses, thereby masking execution-level failure modes under aggregate accuracy metrics. This work proposes a two-tiered evaluation framework that disentangles scorer-agnostic execution states—such as termination, answer exposure, parseability, and output length—from scorer-dependent correctness judgments. By enforcing a fixed output budget, tracking multidimensional execution trajectories, formulate a verification mechanism driven by coverage auditing, the study systematically uncovers divergent execution behaviors across models on MATH and ARC-Challenge benchmarks. The analysis reveals that extended output lengths can mitigate certain failure modes and demonstrates that verification strategies substantially influence comparative accuracy outcomes.
This work addresses the challenge of achieving precise, targeted data removal from the persistent memory of large language models. To accommodate diverse memory representations, the authors propose a "subtract–replay" dichotomous framework: algebraically decomposable memories are updated via subtraction, while entangled representations are reconstructed through deterministic replay. This approach achieves bit-level exact unlearning for the first time in billion-scale models—including Gemma 3 (1B–12B) and the 48B Kimi Linear model—demonstrating that post-deletion model outputs exhibit no statistically significant differences from those of models never exposed to the target data, as measured by KL divergence (as low as 5.4×10⁻¹⁵), perplexity, and robustness against LiRA membership inference attacks, thereby validating both the efficacy and security of the deletion mechanism.
To address unexpected behavioral shifts in language models (LMs) following fine-tuning or deployment, this paper introduces Behavioral Shift Auditing (BSA), a continuous monitoring framework. BSA operates without access to model parameters or gradients, and—uniquely—establishes the first unsupervised, statistical hypothesis testing framework for text generation comparison, leveraging the Kolmogorov–Smirnov test and bootstrap resampling to reliably detect distributional shifts in critical capabilities such as toxicity and translation. The method provides theoretically grounded false positive control and supports configurable tolerance thresholds to accommodate diverse application scenarios. Experiments demonstrate that BSA achieves stable detection of significant behavioral shifts using only hundreds of samples, attaining high sensitivity and low false positive rates on both toxicity and machine translation tasks. Overall, BSA establishes a lightweight, robust, and interpretable paradigm for continuous auditing of LM behavioral evolution.
This paper addresses the lack of standardized, rigorous evaluation benchmarks for large language models (LLMs) in adversarial cybersecurity applications—particularly penetration testing. We systematically review 15 prototype systems and their empirical testing practices drawn from 16 scholarly works. Through comparative analysis of testbed architectures, evaluation metric frameworks, and security experimentation methodologies, we identify three pervasive deficiencies: poor reproducibility, weak real-world representativeness (especially the limited transferability of CTF-based scenarios to actual red-teaming), and narrow assessment scope (lacking established baselines, qualitative analysis, and human-factor considerations). To address these gaps, we propose the first dedicated evaluation framework for LLM-driven adversarial security, featuring an extensible testbed, standardized baseline construction, a hybrid quantitative–qualitative metric suite, and human-in-the-loop assessment dimensions. Based on this framework, we derive seven actionable best-practice recommendations—establishing a scientifically grounded, robust, and operationally viable evaluation paradigm for the field.
This study investigates whether quantization can effectively mitigate the privacy risks posed by large language models’ verbatim reproduction of training data. Departing from conventional membership inference attacks, this work introduces verbatim extraction rate as a novel metric to systematically assess privacy leakage and evaluates the trade-off between memory retention and model performance under various quantization configurations. Using three sizes of the Pythia model family, two quantization algorithms, five precision levels (including 4-bit), and two evaluation corpora, the authors concurrently measure perplexity and verbatim extraction rate. Results show that while quantization reduces memorization more rapidly than it degrades performance, 4-bit quantization in the largest model still retains substantial amounts of training data, highlighting the limitations of quantization as a privacy-preserving technique and revealing its selective forgetting behavior toward memorized content.
This study addresses the confounding influence of probe construction artifacts on self-feedback analyses in language models by proposing a dual-ring co-stochastic evolution framework grounded in cyclic token-cell architectures and Glauber dynamics. Integrating maximal coupling, damage propagation analysis, Lyapunov exponent estimation, and controlled experiments, the approach effectively disentangles spurious signals arising from probe design from genuine model capabilities. The findings reveal that certain phenomena previously attributed to models are in fact artifacts of the probing methodology, identify a model-agnostic damage light cone, and uncover a model-dependent zero-crossing phase transition in Lyapunov exponents. Experimental validation across 19 models and two scale series demonstrates the model-invariance of token-space Lyapunov exponents, accurately reproduces Domany–Kinzel damage fields with high fidelity, and corrects four prior misinterpretations in the literature.
Existing diagnostic methods for assessing memorization in large code models struggle to disentangle memorization from representational capacity as model scale increases, leading to distorted evaluations. This work proposes a novel paradigm that decouples representational load from memorization behavior through invertible mathematical transformations, complemented by systematic analyses employing synonym obfuscation, dead code insertion, and log-probability probing techniques. The study reveals that current probing approaches significantly fail on large-scale models, while the models themselves exhibit strong robustness to diverse surface forms of code. These findings challenge the validity of relying solely on contaminated benchmarks to evaluate memorization and provide both theoretical grounding and methodological support for reassessing the generalization capabilities of large code models.
This work addresses the lack of transparency in current large language model (LLM) services, where users cannot easily verify whether the deployed model matches its claimed identity. Existing identification methods typically require long inputs, fine-grained outputs, or cooperation from model providers. In contrast, this study proposes a lightweight authentication approach that constructs behavioral fingerprints using the output distribution of just a single token elicited by minimal prompts—such as “say a random number between 1 and 100.” The method demonstrates for the first time that single-token distributions alone suffice to reliably distinguish among LLMs without complex inputs or internal access. Evaluated across 165 commercial models using multilingual single-token queries, empirical distribution modeling, and Jensen–Shannon divergence, it achieves intra-model distances an order of magnitude smaller than inter-model distances, 59.5% accuracy in model family identification, and an equal error rate as low as 7.3%, requiring only approximately 100 queries.