evaluate exact recall

Design, implement, and analyze evaluation procedures that measure a model's ability to reproduce exact token sequences (exact recall), including greedy decoding tests, behavioral secret-extraction experiments, and hit@k-style metrics that check whether a target k-token sequence appears among a model's top outputs; also develop protocols to compute and report hit@k and to separate behavioral output-based recall from probabilistic probes such as NLL.

evaluateexactrecall

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.63
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses a critical flaw in existing security evaluations, wherein stored labels are erroneously treated as ground-truth behavioral facts, leading to processing leaks and construct validity threats. To rectify this, the authors propose a seven-stage integrity chain and an endpoint integrity checker that jointly audit historical label assignment mechanisms through execution trace tracing, blind processing reconstruction, double-blind consistency review, and semantic boundary analysis. This approach exposes the non-terminal nature of labels and reconstructs the measurement foundation of security assessments. Empirically, the method corrects 58 mislabeled instances, eliminates all false-positive attack records in the v2 census while preserving genuine privilege escalation cases, and achieves consensus across reviewers on 96 explainable requests.

agent behaviorconstruct validitylabeling bias

Traditional language model evaluation often conflates the ability to produce assessable responses with the correctness of those responses, thereby masking execution-level failure modes under aggregate accuracy metrics. This work proposes a two-tiered evaluation framework that disentangles scorer-agnostic execution states—such as termination, answer exposure, parseability, and output length—from scorer-dependent correctness judgments. By enforcing a fixed output budget, tracking multidimensional execution trajectories, formulate a verification mechanism driven by coverage auditing, the study systematically uncovers divergent execution behaviors across models on MATH and ARC-Challenge benchmarks. The analysis reveals that extended output lengths can mitigate certain failure modes and demonstrates that verification strategies substantially influence comparative accuracy outcomes.

accuracy metricsbenchmarkingexecution outcomes

This work addresses the challenge of achieving precise, targeted data removal from the persistent memory of large language models. To accommodate diverse memory representations, the authors propose a "subtract–replay" dichotomous framework: algebraically decomposable memories are updated via subtraction, while entangled representations are reconstructed through deterministic replay. This approach achieves bit-level exact unlearning for the first time in billion-scale models—including Gemma 3 (1B–12B) and the 48B Kimi Linear model—demonstrating that post-deletion model outputs exhibit no statistically significant differences from those of models never exposed to the target data, as measured by KL divergence (as low as 5.4×10⁻¹⁵), perplexity, and robustness against LiRA membership inference attacks, thereby validating both the efficacy and security of the deletion mechanism.

data removalexact deletionlanguage-model memory

An Auditing Test To Detect Behavioral Shift in Language Models

Oct 25, 2024
LR
Leo Richter
🏛️ University College London | University of Edinburgh | Miniml.AI

To address unexpected behavioral shifts in language models (LMs) following fine-tuning or deployment, this paper introduces Behavioral Shift Auditing (BSA), a continuous monitoring framework. BSA operates without access to model parameters or gradients, and—uniquely—establishes the first unsupervised, statistical hypothesis testing framework for text generation comparison, leveraging the Kolmogorov–Smirnov test and bootstrap resampling to reliably detect distributional shifts in critical capabilities such as toxicity and translation. The method provides theoretically grounded false positive control and supports configurable tolerance thresholds to accommodate diverse application scenarios. Experiments demonstrate that BSA achieves stable detection of significant behavioral shifts using only hundreds of samples, attaining high sensitivity and low false positive rates on both toxicity and machine translation tasks. Overall, BSA establishes a lightweight, robust, and interpretable paradigm for continuous auditing of LM behavioral evolution.

Detect unintended behavioral shifts in language models post-deploymentMonitor changes in model outputs like toxicity and translationProvide a configurable auditing test with theoretical guarantees

This paper addresses the lack of standardized, rigorous evaluation benchmarks for large language models (LLMs) in adversarial cybersecurity applications—particularly penetration testing. We systematically review 15 prototype systems and their empirical testing practices drawn from 16 scholarly works. Through comparative analysis of testbed architectures, evaluation metric frameworks, and security experimentation methodologies, we identify three pervasive deficiencies: poor reproducibility, weak real-world representativeness (especially the limited transferability of CTF-based scenarios to actual red-teaming), and narrow assessment scope (lacking established baselines, qualitative analysis, and human-factor considerations). To address these gaps, we propose the first dedicated evaluation framework for LLM-driven adversarial security, featuring an extensible testbed, standardized baseline construction, a hybrid quantitative–qualitative metric suite, and human-in-the-loop assessment dimensions. Based on this framework, we derive seven actionable best-practice recommendations—establishing a scientifically grounded, robust, and operationally viable evaluation paradigm for the field.

Assessing benchmarking practices in LLM-based penetration testingBridging gaps between security research and real-world practiceEvaluating LLM-driven offensive cybersecurity tools' effectiveness

Latest Papers

What's happening recently
View more

This study investigates whether quantization can effectively mitigate the privacy risks posed by large language models’ verbatim reproduction of training data. Departing from conventional membership inference attacks, this work introduces verbatim extraction rate as a novel metric to systematically assess privacy leakage and evaluates the trade-off between memory retention and model performance under various quantization configurations. Using three sizes of the Pythia model family, two quantization algorithms, five precision levels (including 4-bit), and two evaluation corpora, the authors concurrently measure perplexity and verbatim extraction rate. Results show that while quantization reduces memorization more rapidly than it degrades performance, 4-bit quantization in the largest model still retains substantial amounts of training data, highlighting the limitations of quantization as a privacy-preserving technique and revealing its selective forgetting behavior toward memorized content.

language modelsLLM quantizationprivacy risk

This study addresses the confounding influence of probe construction artifacts on self-feedback analyses in language models by proposing a dual-ring co-stochastic evolution framework grounded in cyclic token-cell architectures and Glauber dynamics. Integrating maximal coupling, damage propagation analysis, Lyapunov exponent estimation, and controlled experiments, the approach effectively disentangles spurious signals arising from probe design from genuine model capabilities. The findings reveal that certain phenomena previously attributed to models are in fact artifacts of the probing methodology, identify a model-agnostic damage light cone, and uncover a model-dependent zero-crossing phase transition in Lyapunov exponents. Experimental validation across 19 models and two scale series demonstrates the model-invariance of token-space Lyapunov exponents, accurately reproduces Domany–Kinzel damage fields with high fidelity, and corrects four prior misinterpretations in the literature.

damage spreadinglanguage modelsmeasurement confounding

Existing diagnostic methods for assessing memorization in large code models struggle to disentangle memorization from representational capacity as model scale increases, leading to distorted evaluations. This work proposes a novel paradigm that decouples representational load from memorization behavior through invertible mathematical transformations, complemented by systematic analyses employing synonym obfuscation, dead code insertion, and log-probability probing techniques. The study reveals that current probing approaches significantly fail on large-scale models, while the models themselves exhibit strong robustness to diverse surface forms of code. These findings challenge the validity of relying solely on contaminated benchmarks to evaluate memorization and provide both theoretical grounding and methodological support for reassessing the generalization capabilities of large code models.

code LLMsmemorizationprobing techniques

This work addresses the lack of transparency in current large language model (LLM) services, where users cannot easily verify whether the deployed model matches its claimed identity. Existing identification methods typically require long inputs, fine-grained outputs, or cooperation from model providers. In contrast, this study proposes a lightweight authentication approach that constructs behavioral fingerprints using the output distribution of just a single token elicited by minimal prompts—such as “say a random number between 1 and 100.” The method demonstrates for the first time that single-token distributions alone suffice to reliably distinguish among LLMs without complex inputs or internal access. Evaluated across 165 commercial models using multilingual single-token queries, empirical distribution modeling, and Jensen–Shannon divergence, it achieves intra-model distances an order of magnitude smaller than inter-model distances, 59.5% accuracy in model family identification, and an equal error rate as low as 7.3%, requiring only approximately 100 queries.

behavioral fingerprintinglarge language modelsmodel identification

Hot Scholars

YL

Yixuan Liu

AMD, Tsinghua University
Generative AI
ZX

Ziqi Xu

Lecturer, School of Computing Technologies, RMIT University
Causal AIFairness
YF

Ying Fang

Westlake University; Zhejiang University
speech recognition
KY

Kun Yuan

University of Strasbourg & Technical University of Munich
surgical data sciencemulti-modal learning
SR

Simon Razniewski

Professor at ScaDS.AI & TU Dresden
Language ModelsKnowledge BasesCommonsense KnowledgeNLP