Score
Designs and implements pre-registered evaluation protocols, metrics, and analysis plans to quantify a system's empirical robustness to noise, corruptions, and other perturbations; this includes specifying perturbation sets, comparative degradation hypotheses, evaluation metrics (accuracy and robustness metrics), and pre-specified statistical analyses. Builds reproducible assessment pipelines and reporting procedures that apply production-relevant perturbations, measure performance across conditions and controllers, and disclose outcomes according to the pre-registration.
研究通过审计Praxa AI管道文件及记录,发现代理评估指标与标签含义不一致的问题,并提供了一个可重用的验证包来区分不同类型的性能声明。
To address the longstanding trade-off between computational cost and accuracy in pre-deployment robustness assessment for safety-critical applications, this paper proposes a hypothesis-testing-based quantitative evaluation framework. Our core contribution is the introduction of “tower robustness”—a novel metric that, for the first time, incorporates statistical hypothesis testing into probabilistic modeling of deep learning robustness, enabling rigorous, verifiable quantification of model output stability under input perturbations. By integrating probabilistic modeling with comparative analysis, the framework systematically restructures the evaluation pipeline, achieving both theoretical soundness and substantial efficiency gains. Extensive experiments across large-scale benchmarks demonstrate that our approach improves assessment accuracy by 12.7% on average and reduces runtime by 43.5% compared to state-of-the-art baselines. This work establishes a new paradigm for pre-deployment risk analysis of high-assurance AI systems—one that is both practically deployable and inherently interpretable.
This study addresses the reliability of probabilistic uncertainty quantification (UQ) in software defect prediction, particularly its ability to reflect model performance and calibration—especially in cross-project settings, where systematic validation remains lacking. Through a large-scale empirical analysis of 16 classifiers across 36 within-project and 32 cross-project datasets, the work examines the relationships between five UQ metrics and six performance measures alongside three calibration metrics. It reveals, for the first time, a strong context dependency: within-project, UQ correlates strongly with false positive rate and AUC, but these correlations substantially weaken or even reverse in cross-project scenarios. Notably, high-performing models can still exhibit severe miscalibration. These findings indicate that UQ signals are not directly transferable and must be evaluated independently relative to specific objectives, using multidimensional calibration assessments.
This work addresses the efficiency bottleneck of manual reproducibility reviews in safety-critical domains such as the Internet of Things and cyber-physical systems, which hampers research transparency and deployability. The paper presents the first systematic framework leveraging large language models (LLMs) to automate reproducibility assessment by integrating natural language understanding, code generation, sandboxed environment auto-configuration, and rule-guided flaw detection. This approach enables reproducibility scoring, automatic execution environment setup, and identification of methodological flaws. Experimental results demonstrate that the proposed method achieves over 72% accuracy in reproducibility judgment, automatically constructs executable environments for 28% of runnable artifacts, and attains F1 scores exceeding 92% across seven common categories of methodological defects, substantially enhancing both the efficiency and quality of reproducibility review.
In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.
Current AI-assisted scientific writing lacks auditable generation processes and mechanisms for accountability, undermining the verifiability of research credibility and compliance. This work proposes a novel auditing paradigm embedded directly within the production workflow, enforcing end-to-end traceability, immutability, and third-party reproducibility of AI involvement through preregistered blind-spot indicator cards, sealed execution environments, and automated gatekeeping intercepts. Core technical components include Git-sealed lineage anchoring, hash-bound provenance tracking, red-flag interception protocols, cross-model role isolation, and programmatic assembly. In experimental validation, one project was automatically terminated when preregistered confirmatory tests triggered a No-Go decision. An open-source toolkit is released to enable independent recomputation of all core audit metrics by third parties.
研究解决了语言模型评估不稳定的问题,通过预注册审计发现请求重复性和一致性未达预期标准,提出设计规则和报告清单以改善测量可靠性。
This work addresses a critical gap in the evaluation of security vulnerability reproductions generated by large language models (LLMs) or agents: while such artifacts are often executable, they frequently lack rigorous validation confirming that they faithfully reproduce the specific CVE in question, revealing a disconnect between semantic intent and actual vulnerability signals. To remedy this, the paper introduces the first reusable evaluation protocol for security reproductions, integrating preregistered auditing, an R0/R1 environment remediation ladder, G1–G3 semantic evidence tiers, and a patch-based counterfactual oracle. Applying this framework to 104 published artifacts from 2023–2026, the study finds that 56.9% exhibit CVE identifier mismatches, only 61.1% remain functional after patching, and the oracle demonstrates limited reliability with 60% sensitivity and 45% specificity—highlighting widespread issues of false-positive patch versions and benign inputs in current approaches.
本文研究了工业控制系统中异常检测模型在训练数据受污染情况下的鲁棒性,通过三种不同的数据污染策略评估了11种不同检测器的表现。
本文研究多阶段LLM管道中证据状态的可靠性问题,通过引入Evidence-State Reliability (ESR)评估方法,在不同证据条件下测试了GLM-5.2模型的表现。