Score
Designs, builds, and evaluates methods, tools, and procedures to detect, analyze, and quantify intentional or unintentional contamination of datasets and data pipelines (e.g., poisoned examples or silent corruption), including algorithms for contamination detection, forensic analysis, and impact assessment. Also develops and operationalizes mitigation workflows and controls—such as removal, quarantining, retraining, validation checks, and monitoring—to reduce contamination effects and prevent recurrence.
This work investigates how external uncertainties propagate through structured multi-agent workflows to induce information contamination, thereby degrading reasoning trajectories and output correctness. We introduce a taxonomy of three distinct manifestations of information contamination along with their control-flow characteristics, establishing the first classification framework tailored to structured multi-agent workflows and a trajectory-based detection and localization methodology. Through systematic injection of structured perturbations across 32 GAIA tasks and 614 experimental configurations involving three diverse models, we uncover a decoupling between workflow structural divergence and answer correctness, exposing the fundamental limitations of current validation mechanisms. These findings provide empirical grounding for the design of robust, defense-oriented multi-agent workflows.
In highly regulated domains such as finance, data quality control (QC) is often fragmented into isolated preprocessing steps, undermining end-to-end trustworthy AI pipelines. To address this, we propose the first AI-driven DataOps framework that embeds QC as a system-level core component. Our framework deeply integrates rule-based engines, statistical analysis, and custom AI-powered anomaly detection across the entire data lifecycle—from ingestion and transformation to model deployment—enabling dynamic remediation, policy-configurable workflows, and end-to-end auditability. Technically, it unifies data profiling, stream processing, cloud-native storage interfaces, and a proprietary AI detection module. Evaluated in a real-world financial production environment, the framework achieves significantly improved anomaly recall, reduces manual intervention by 42%, and ensures audit completeness and full data traceability under high-throughput conditions, fully satisfying regulatory compliance requirements.
This work addresses the pervasive issue of redundant and isolated messages in system logs, which hinder downstream tasks such as model reasoning and anomaly detection. To tackle this challenge, the authors propose LogPurifier—the first task-agnostic log cleansing framework—that systematically purifies logs by extracting log templates and modeling their dependencies to accurately identify and remove messages irrelevant to system functional behavior. By doing so, LogPurifier enables effective log sanitization applicable across diverse analytical scenarios. Experimental results demonstrate that LogPurifier substantially improves both accuracy and efficiency in various downstream tasks, thereby validating its effectiveness and generalizability.
Current automated detection tools struggle to meet regulatory practice demands due to insufficient transparency, interpretability, and the inability to map findings to specific legal provisions, resulting in a disconnect between academic research and enforcement applications. Through in-depth interviews with nine regulatory practitioners and an analysis integrating regulatory workflows with technical feasibility, this study systematically uncovers, from a regulatory perspective, the practical barriers to deploying automated tools for identifying deceptive designs. The work proposes a human-in-the-loop compliance review framework that is user-need-driven, supports the entire investigative workflow, and aligns both research and regulatory objectives, offering critical guidance for developing automated detection systems that genuinely meet real-world enforcement requirements.
This study addresses the limitations of existing defect detection methods in bioinformatics, which are hindered by the absence of high-quality, domain-specific datasets. To bridge this gap, we introduce BioDefect, the first defect detection dataset tailored for bioinformatics software, constructed from real-world code repositories with full contextual preservation and rigorous mitigation of label inconsistency and data leakage. We systematically evaluate BioDefect on nine language models, including DeepSeek-R1, demonstrating substantial performance gains: all models achieve an average F1-score improvement ranging from 29.61% to 38.04% over those trained on existing general-purpose datasets. These results underscore the effectiveness and superiority of BioDefect in advancing defect detection within the bioinformatics domain.
This study addresses the challenge of inflated performance of large language models (LLMs) in software engineering evaluations due to pretraining data contamination, a problem exacerbated by the opacity of their training corpora. The work formalizes this issue as a white-box membership inference task over source code and introduces a general-purpose detection framework that operates across models and datasets. To catalyze progress in this area, the authors organized the first “Poisoned Cup” LLM evaluation competition (FSE-AIWare 2026), providing a standardized dataset, target models, and baseline methods. This initiative advances the development of generalizable decontamination-aware evaluation techniques and fosters a more trustworthy ecosystem for assessing LLMs in software engineering contexts.