Score
Plans and performs systematic audits of systems, software artifacts, and operational configurations to detect security weaknesses, compromises, misconfigurations, or noncompliance. This work includes testing for poisoned artifacts and bundles, verifying certificate or equivalence-chain validity, assessing resource and access requirements and policy/legal implications, and producing documented findings with prioritized mitigations.
SBOM-driven vulnerability scanning (SVS) tools suffer from inconsistent results and silent failures—causing false positives or negatives—that undermine security assurance. To address this, we propose SVS-TEST, the first systematic evaluation framework for SVS tools, comprising a rigorous methodology, an open-source toolchain, and a benchmark suite featuring 16 carefully constructed SBOM samples paired with ground-truth vulnerability annotations. SVS-TEST enables reproducible, quantitative assessment of SVS tools’ capabilities, maturity, and error-handling behaviors via automated orchestration, cross-tool discrepancy analysis, and root-cause attribution of failures. Empirical evaluation across seven widely adopted SVS tools reveals substantial reliability disparities: several tools fail silently on syntactically and semantically valid SBOMs. All artifacts—including the benchmark, toolchain, and evaluation reports—have been open-sourced, and findings were responsibly disclosed to affected tool maintainers prior to publication.
This work addresses the critical yet often overlooked security risks in large language model (LLM) agent systems, which frequently stem from software stack components such as tool code, deployment configurations, and permission settings—not merely from the underlying models. To this end, we present the first dedicated security analysis framework tailored for LLM-based agent applications. Our system integrates dataflow analysis, credential detection, structured configuration parsing, and permission risk assessment to precisely identify diverse vulnerabilities across tool functions, prompts, and deployment artifacts, outputting results in the standardized SARIF format. Evaluated on 22 real-world samples containing 42 annotated vulnerabilities, our approach successfully detects 40 true positives with only 6 false positives, substantially outperforming general-purpose static application security testing (SAST) tools while completing each scan in under one second.
Traditional code auditing tools struggle to identify security vulnerabilities implied in natural language specifications and often produce false positives with ambiguous root causes. This work proposes SPECA, a framework that parses natural language specifications to extract explicit, typed security properties and integrates formal modeling with structured proof-based reasoning for vulnerability detection. SPECA’s key innovations include specification-aware precise auditing, unified cross-repository property comparison, and a traceable false positive attribution mechanism. Experimental evaluation demonstrates that SPECA successfully reproduces all known vulnerabilities on the Sherlock and RepoAudit benchmarks while uncovering multiple previously unknown flaws; furthermore, its false positives are systematically attributable to three distinct, actionable root causes amenable to improvement.
This work addresses the challenge that ambiguous requirements in natural language specifications often lead to consistent errors across multiple implementations, which traditional differential testing fails to detect. To this end, the authors propose SPECA, a novel framework that automatically translates informal specifications into structured checklists and maps them to critical code locations across diverse implementations, enabling checklist-driven, one-to-many cross-implementation auditing. SPECA integrates natural language processing, threat modeling, and agent-based automated auditing to construct an end-to-end specification alignment verification system. Evaluated on the Ethereum Fusaka upgrade, the approach identified 76.5% of valid vulnerabilities through cross-implementation checks, and its optimized auditing agent achieved a 27.3% recall rate on high-severity vulnerabilities, outperforming 96% of human auditors.
To address the challenges of identifying static security vulnerabilities in proprietary and open-source software, unclear vulnerability remediation priorities, and escalating software supply chain risks, this paper proposes an end-to-end, customizable Static Application Security Testing (SAST) workflow. The workflow enables multi-tool orchestration, iterative scanning, and seamless DevSecOps integration, incorporating AI-driven vulnerability prioritization and automated remediation governance as key innovations. Leveraging a generalized process design with environment-adaptive configuration, it significantly improves detection coverage and remediation efficiency. Experimental evaluation in industrial settings demonstrates that the approach reduces source-code-level vulnerabilities by 32.7%, mitigates third-party component–introduced risks by 41.5%, and ensures backward compatibility with legacy systems while supporting scalable deployment across heterogeneous environments.
This work addresses the lack of systematic evaluation benchmarks for large language models (LLMs) in security audit log investigation tasks by introducing AuditBench, the first audit log benchmark specifically designed for attack investigation. AuditBench encompasses over 50 real-world scenarios across Linux and Windows systems and focuses on four core tasks: alert classification, persistence mechanism identification, among others. Through multidimensional experiments, the study systematically evaluates the impact of model scale, log representation, prompt design, and fine-tuning strategies on performance and error patterns, while also analyzing the quality of LLM-generated explanations. The findings reveal the capability boundaries and characteristic failure modes of various models across different investigative tasks, providing empirical foundations for deploying and optimizing LLMs in security operations.
This study addresses the limitations of current AI auditing practices, which predominantly focus on individual models while overlooking integration risks arising from interactions among system components and between systems and their environments. Through a scoping review and reflexive thematic analysis of 58 studies, the work systematically codes existing literature to delineate, for the first time, three distinct domains of AI integration auditing: inter-component, system–environment, and multi-system. It further introduces domain-specific evaluation dimensions—compatibility, completeness, and oversight—that capture unique aspects of integrated AI systems. The findings reveal that current auditing practices remain fragmented and nascent, underscoring the critical role of accessible information and resource support in effective audit design. The paper calls for novel auditing frameworks capable of spanning components, environments, and systems to enable systematic exploration, identification, coordination, and standardization of integration-related risks.
This study addresses the prevalence of logical flaws in business process documentation—often stemming from conflicting requirements, ambiguous phrasing, and insufficient quality assurance—which frequently lead to product defects, project delays, and cost overruns. It presents the first systematic exploration of end-to-end applications of large language models (LLMs) in this domain: automatically extracting business logic from unstructured sources such as ISO standards and user manuals, constructing attributed logic graphs, and integrating graph-based analysis with formal verification techniques to detect latent vulnerabilities. Empirical evaluation demonstrates the effectiveness of multiple LLMs across critical tasks including grammatical error correction, identification of technical inaccuracies, and reconstruction of process structures, substantially advancing the automation of flaw detection in business process specifications.
This work addresses the vulnerability of enterprise software supply chains to infrastructure-level attacks and the limitations of traditional verification approaches that rely on consumers re-executing builds and tests—a process that incurs significant bottlenecks and erodes trust. To overcome these challenges, the paper proposes an evidence-driven, trustworthy CI pipeline protocol that, for the first time, integrates deterministic builds (based on Nix) with remote attestation from Intel TDX trusted execution environments (TEEs). The authors formally define the evidence lifecycle and introduce lightweight cryptographic signing and policy validation mechanisms. This approach provides strong cryptographic guarantees of integrity, authenticity, and provability for CI artifacts without requiring consumer-side re-execution, substantially improving verification efficiency and system scalability while effectively amortizing the initial overhead introduced by TEEs.