threshold and error-rate analysis

Computing and evaluating operating thresholds and error-rate metrics (e.g., equal error rate, word error rate) and their effects on detection, privacy, and performance under attacker models. Employed to set unsupervised thresholds, compare to labeled baselines, and quantify failure modes that impact metric values.

thresholdanderror-rateanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the widespread issue in AI red-teaming evaluations where comparisons of attack success rates (ASR) often rely on invalid or incomparable measurements, leading to erroneous conclusions about system security or attack efficacy. For the first time, the paper introduces measurement validity theory from social sciences into the AI red-teaming domain, systematically analyzing—through inferential statistics and illustrative case studies such as jailbreaking attacks—the conditions under which ASR comparisons are meaningful. The work establishes clear prerequisites for valid ASR comparisons, identifies and categorizes common patterns of invalid comparison, and thereby provides a rigorous theoretical foundation that significantly enhances the scientific rigor and comparability of AI safety evaluations.

AI red teamingattack success ratecomparability

Traditional automatic speech recognition evaluation metrics, such as word error rate (WER) and character error rate (CER), fail to capture human perception of errors and neglect linguistic and semantic influences. This work proposes a novel paradigm that embeds any perception-oriented evaluation metric into the minimum edit distance (minED) framework to produce an intuitively interpretable equivalent error rate. For the first time, this approach translates human perceptual modeling into a comprehensible error rate format, enabling quantification of error severity from the perspective of human understanding. The resulting metric not only aligns closely with human judgments but also effectively identifies recognition errors that critically impact semantic comprehension.

Automatic Speech RecognitionCharacter Error Rateevaluation metrics

Traditional vulnerability scoring systems (e.g., CVSS, EPSS) exhibit fundamental limitations in assessing adversarial attacks (AAs) against large language models (LLMs), as their static, context-agnostic design fails to capture semantic-level threat distinctions—particularly for prompt injection and other input-manipulation attacks. Method: We systematically evaluate 56 adversarial attack variants across GPT-4, Claude, and Llama, quantifying score dispersion via the coefficient of variation (CV) and benchmarking against multi-source adversarial datasets. Contribution/Results: Empirical analysis reveals that existing metrics assign uniformly low, highly convergent scores to most attacks (CV < 0.08), demonstrating severe discriminative inadequacy. To address this, we propose the first flexible, context-aware scoring framework specifically designed for generative AI—grounded in adaptive risk modeling, dynamic contextualization, and semantic impact quantification. This work establishes a novel paradigm for rigorous, actionable LLM security assessment.

Adversarial AttacksLarge Language ModelsVulnerability Assessment

This study addresses the challenge posed by severe class imbalance in anomaly detection, which complicates the interpretation and comparison of common evaluation metrics. The authors systematically analyze the behavior of AUROC, AUPR, F1-score, and Matthews Correlation Coefficient (MCC) across varying anomaly ratios and introduce a novel "metric landscape" visualization technique. This approach reveals, for the first time, each metric’s inherent preference for true positive rate versus true negative rate and how their stability varies with imbalance levels. By modeling the relationship between metrics and anomaly prevalence, the work delineates clear applicability boundaries for each metric, thereby providing a principled, interpretable foundation for reliable metric selection in highly imbalanced anomaly detection scenarios.

anomaly detectionclass imbalanceevaluation metrics

Latest Papers

What's happening recently
View more

This work addresses the risk that online platforms may strategically generate semantically equivalent content variants to manipulate compliance metrics, creating a “gaming” problem where apparent metric improvements mask unmitigated harms. The authors model moderation protocols as transformation graphs and introduce a semantic envelope metric, theoretically proving it to be the pointwise minimal solution within the class of conservative repairs. They further develop a hierarchical certification mechanism that guarantees effective constraint of true harm under any policy. Experimental evaluation—combining finite-state mixed-strategy enumeration, SMT solving (using Z3 and cvc5), and bounded single-player MDP verification in PRISM-games—demonstrates that conventional metrics often exhibit significant violations and gaming gaps, whereas the semantic envelope metric remains violation-free across all test instances, effectively resisting strategic manipulation.

audit certificationmetric manipulationonline safety regulation

This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.

calibrationexchangeabilityfalse sense of safety

Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.

class imbalancemodel evaluationperformance metrics

This work addresses the critical issue of evaluator instability in red-teaming evaluations of large language models (LLMs), where attack success rates (ASR) are highly sensitive to the choice of evaluator, leading to unreliable assessments. To mitigate this, the authors propose a two-stage, reliability-aware evaluation framework: first identifying high-disagreement attack categories through multi-evaluator comparison and uncertainty quantification, then selectively refining ASR estimates via an independent, annotation-free validation mechanism. Experiments on Garak reveal significant evaluator disagreement across 22 out of 25 attack categories, with ASR varying by up to 33% depending on the evaluator. The proposed method improves evaluation accuracy from 72% to 89%, offering the first systematic quantification of evaluator instability and enabling more reliable and controllable ASR estimation.

attack success rateevaluator instabilityLLM red-teaming

Existing prompt injection detectors frequently miss high-severity attacks under distribution shifts while still outputting overconfident predictions near 1.0, a risk substantially underestimated by standard calibration metrics. This work introduces a severity-aware calibration perspective, proposing a severity metric S to quantify the confidence assigned to missed attacks. Through fixed-threshold cross-distribution evaluation, black-box rewrite-based attack generation, and instruction-tuned models as judges, the study systematically investigates detector failure mechanisms. Findings reveal that content keywords—not injection syntax—are the primary cause of detection blind spots. Critically, all evaluated detectors exhibit highly confident false negatives across diverse distribution shifts, exposing fundamental limitations in current calibration approaches.

attack shiftcalibrationconfidence

Hot Scholars

GC

Giuseppe Caire

Professor, Technical University of Berlin, Germany, and Professor of Electrical Engineering (on
Information TheoryCommunicationsSignal ProcessingStatistics
WZ

Wenjun Zhang

City University of Hong Kong
Thin film technologynanomaterials and nanodevices
LS

Laurent Schmalen

Professor | Fellow IEEE | Communications Engineering Lab, Karlsruhe Institute of Technology
Information TheoryCoding TheoryError Correction CodingOptical Communications
GG

Gennian Ge

Capital Normal University
CombinatoricsCoding theoryInformation Security
BS

B. Sundar Rajan

Electrical Communication Engineering Department, Indian Institute of Science
Wireless CommunicationCoding TheoryInformation TheoryNetwork Coding