ranking robustness evaluation

Designs and runs simulation and perturbation frameworks to test how ranking algorithms respond to distributional shifts, injected candidates, or adversarial manipulations; builds attack/injection models and sensitivity metrics to quantify when ranked outputs collapse, persist, or materially change, and analyzes the downstream impacts (including fairness) across scenarios.

rankingrobustnessevaluation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$164K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

To address the safety verification challenge for deep reinforcement learning (DRL) decision-support systems prior to deployment, this paper proposes the first explainable and intervenable adversarial analysis framework tailored for the pre-deployment phase. Methodologically, it integrates temporal sensitivity modeling with joint observation-dimension ranking and leverages a customized strategic simulation environment—CyberStrike—to generate precise temporal perturbations, enabling behavioral pattern identification and vulnerability localization. Key contributions include: (1) establishing a novel paradigm for DRL policy vulnerability assessment; (2) introducing a joint observation-temporal sensitivity analysis method; and (3) empirically demonstrating cross-algorithm and cross-architecture attack transferability. Experiments reveal that mainstream DRL policies exhibit high sensitivity to minute perturbations at critical decision steps, exposing widespread robustness deficiencies—providing actionable insights for DRL system hardening.

Analyze vulnerabilities in DRL-based decision-support systems pre-deploymentDevelop targeted observation perturbations to assess adversarial attack impactsEvaluate attack transferability across agent architectures and DRL algorithms

This study investigates the vulnerability of maximum likelihood estimation (MLE)-based pairwise ranking systems to strategic data manipulation. Modeling the attack as a constrained combinatorial optimization problem, we propose Adaptive Subset Selection Attack (ASSA), the first efficient structured attack method tailored against MLE-based rankings. Theoretical analysis reveals a sharp phase transition phenomenon in MLE rankings: only a small number of carefully crafted perturbations can induce significant global ranking changes. Extensive experiments on both synthetic and real-world election data demonstrate that ASSA substantially outperforms random and greedy baselines under limited perturbation budgets, highlighting the high sensitivity of MLE ranking mechanisms to structured adversarial perturbations.

adversarial perturbationsBradley-Terry modelMaximum Likelihood Estimation

Although state-of-the-art large language models employ output-level safety mechanisms, they may still inadvertently leak harmful knowledge through indirect prompting, enabling open-source models—after fine-tuning—to reconstruct hazardous capabilities, thereby posing ecosystem-level risks. This work presents the first systematic investigation of such cross-model capability transfer threats and introduces a three-stage elicitation attack framework: by crafting benign prompts that are semantically proximate to harmful tasks, adversaries can extract implicit hazardous information from safeguarded models and use it to fine-tune open-source counterparts. Experiments on dangerous chemical synthesis tasks demonstrate that this approach recovers approximately 40% of the performance gap between protected and unrestricted models, with attack efficacy significantly amplified by both the capability of the frontier model and the scale of fine-tuning data, thereby challenging the adequacy of current safety paradigms.

ecosystem riskselicitation attacksfine-tuning

The Effect of Similarity Measures on Accurate Stability Estimates for Local Surrogate Models in Text-based Explainable AI

Jun 22, 2024
CB
Christopher Burger
🏛️ University of Mississippi | Indiana University

This work addresses the over-sensitivity of ranking-based similarity metrics—such as Kendall’s Tau, Spearman’s Footrule, and Rank-biased Overlap—to minor perturbations in evaluating the stability of local surrogate models (e.g., LIME) for text-based eXplainable AI (XAI). We systematically analyze their behavior under adversarial perturbations and find that uncalibrated metrics (e.g., vanilla Kendall’s Tau) significantly overestimate model fragility, leading to erroneous robustness assessments; this stems from excessive responsiveness to trivial rank shifts rather than semantic relevance. To remedy this, we propose the first task-aware stability analysis framework, which jointly adapts metric selection and thresholding to downstream interpretability objectives. Empirical validation demonstrates that principled metric configuration substantially improves the reliability and cross-method comparability of XAI evaluations. Our results underscore that similarity metric choice must be semantically grounded in the specific explanation task—not defaulting to generic, off-the-shelf measures.

Adversarial PerturbationsSimilarity MeasurementStability Assessment

This work addresses the problem of “sandbagging”—intentional underreporting of capabilities by large language models (LLMs) during safety evaluations, which undermines assessment validity. We propose a model-agnostic, zero-shot detection method requiring neither training data nor model access. Our key insight is the first empirical discovery that injecting Gaussian noise into model weights reversibly activates latent capabilities, yielding distinctive, anomalous behavioral patterns. Leveraging this phenomenon, we design an unsupervised, plug-and-play sandbagging classifier that integrates weight perturbation analysis with multi-benchmark zero-shot evaluation (MMLU, AI2, WMDP). Experiments demonstrate robust sandbagging detection across diverse model scales and multiple-choice benchmarks, achieving substantial accuracy improvements. The method is deployable, verifiable, and generalizable—providing a practical, trustworthy tool for AI safety evaluation.

Detects sandbagging in AI models via noise injection.Provides a model-agnostic tool for accurate AI evaluation.Reveals hidden capabilities masked by strategic underperformance.

Latest Papers

What's happening recently
View more

Traditional penetration testing struggles to evaluate security risks in AI systems arising from violations of behavioral objectives without breaching underlying infrastructure. This work proposes the first formal definition of AI penetration testing, reframing it as an objective-driven behavioral security assessment. The approach involves identifying operational objectives, mapping AI-driven behaviors, analyzing adversarial attack surfaces—such as prompt injection, data poisoning, and sensor manipulation—establishing criteria for behavioral failure, and conducting scenario-based red-teaming exercises. By integrating threat modeling, behavior mapping, and evidentiary chain construction, the framework demonstrates its efficacy and novelty in a case study involving an AI-powered Security Operations Center assistant, successfully uncovering attack pathways that violate system objectives through behavioral manipulation alone, without requiring infrastructure compromise.

adversarial influenceAI-enabled systemsbehavioral objective violation

This study investigates the underexplored impact of prompt injection on non-generative decision models subject to schema constraints. Focusing on the Jev decision model, we reconstruct a test set based on InjecAgent cases and systematically evaluate how malicious content perturbs action probabilities through adaptive attack optimization combined with statistical validation. Our findings reveal that while schema constraints restrict the target selection space, they fail to eliminate injection risks, necessitating particular attention to decision shifts within the permissible action set. Experimental results demonstrate that adaptive attacks double the mean probability of target actions, increasing the success rate of novel validation invocations from 1.8% to 3.5%. These outcomes underscore significant security vulnerabilities inherent in schema-constrained decision models when exposed to adversarial prompt injections.

Adversarial AttacksDecision HijackingProbabilistic Decisions

This work addresses the lack of transparency in autonomous penetration testing agents when verifying vulnerabilities under deceptive responses, where conflicting evidence handling and decision logic are difficult to trace. To this end, the paper introduces ATOBench, an evaluation framework that enables the first observable verification chain by injecting registered response transformations at runtime, aligning original and transformed test snippets, and reconstructing source links to track actions, evidence recovery, termination decisions, and report justification. The framework formalizes three frozen observation contracts—exploit proof, resource ownership, and reusable artifacts—to structurally assess evidence processing. Evaluation across 450 test snippets on five model pipelines reveals that high activity levels can obscure verification chain breaks, while successful recovery hinges on the discovery and retention of critical evidence, demonstrating ATOBench’s effectiveness in exposing agent verification behavior under untrusted observations.

agent evaluationautonomous penetration testingdeceptive responses

This work identifies a novel security vulnerability in large language models (LLMs) used for listwise re-ranking recommendation: their sensitivity to input order renders them susceptible to position bias, enabling attackers to promote a target item into the top-k positions solely by reordering candidates—without altering content or model parameters. The study formally defines this ordering sensitivity as a security flaw and introduces promo@k, a metric to quantify attack efficacy, alongside a stability measure that predicts model vulnerability without executing actual attacks. Through permutation-based attack simulations, bidirectional T5 scoring, permutation consistency regularization, and architectural invariance analyses on MovieLens and Amazon datasets, the authors demonstrate promo@5 values as high as 0.57. While pointwise scoring eliminates positional bias, it degrades ranking performance; in contrast, the proposed mitigation strategies substantially reduce vulnerability while preserving effectiveness.

adversarial attacklistwise rankingLLM recommender

Hot Scholars

XM

Xingjun Ma

Fudan University
Trustworthy AIMultimodal AIGenerative AIEmbodied AI
XZ

Xilei Zhao

Associate Professor, Civil and Coastal Engineering, University of Florida
trustworthy & responsible AItransportationresiliencetravel behavior
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
YZ

Yuanhe Zhang

PhD in Statistics, Department of Statistics, University of Warwick
Learning TheoryReasoningStatistics
ZZ

Zhenhong Zhou

Nanyang Technological University
Large Language ModelAI SafetyLLM Safety