Score
Designs, implements, and maintains coding standards, automated checks, developer tooling, and processes that prevent, detect, and remediate security defects in source code across the software development lifecycle. Performs threat-informed secure code reviews and code deobfuscation, and integrates security-as-code policies and build- and runtime controls into CI/CD and productionization pipelines to ensure secure software engineering and delivery.
This work addresses the vulnerability of build system code to poisoning attacks, which pose a critical threat to software supply chain security. While existing tools primarily focus on application source code, they largely overlook the security of the build process itself. To bridge this gap, we propose a novel paradigm—“development-phase isolation”—that, for the first time, incorporates build scripts into the scope of security analysis. By leveraging information flow tracking and behavioral privilege modeling, our approach enables fine-grained monitoring of build-time code execution. We implement this methodology in a prototype tool, Foreman, which effectively detects anomalous and malicious behaviors within build scripts. In real-world evaluations, Foreman successfully identified the poisoned test files used in the recent XZ Utils supply chain attack, demonstrating both the efficacy and practicality of our approach.
本文针对软件开发过程中早期识别和修复安全漏洞的问题,提出了一种将静态应用安全测试工具输出集成到CI/CD流水线及问题跟踪软件中的自动化方法。
To address the challenges of identifying static security vulnerabilities in proprietary and open-source software, unclear vulnerability remediation priorities, and escalating software supply chain risks, this paper proposes an end-to-end, customizable Static Application Security Testing (SAST) workflow. The workflow enables multi-tool orchestration, iterative scanning, and seamless DevSecOps integration, incorporating AI-driven vulnerability prioritization and automated remediation governance as key innovations. Leveraging a generalized process design with environment-adaptive configuration, it significantly improves detection coverage and remediation efficiency. Experimental evaluation in industrial settings demonstrates that the approach reduces source-code-level vulnerabilities by 32.7%, mitigates third-party component–introduced risks by 41.5%, and ensures backward compatibility with legacy systems while supporting scalable deployment across heterogeneous environments.
研究通过观察100名参与者如何评估和选择AI生成的C语言代码的安全性和功能性,探讨了开发者在使用AI辅助编程时对安全性的关注程度及信任影响。
This study investigates the practical efficacy of GitHub Copilot’s code review capability in detecting security vulnerabilities. Method: We constructed a manually annotated, multilingual dataset of open-source project vulnerabilities—covering SQL injection, cross-site scripting (XSS), and insecure deserialization—and conducted controlled experiments to evaluate Copilot’s feedback along three dimensions: accuracy, vulnerability coverage, and risk-level alignment. Contribution/Results: Empirical results show that Copilot detects fewer than 5% of high-severity vulnerabilities, predominantly flagging low-risk stylistic issues instead. Its security detection performance is substantially inferior to both specialized static application security testing (SAST) tools and human security audits. To our knowledge, this is the first empirical study to expose fundamental limitations of AI-powered programming assistants in security-critical code review tasks. Our findings challenge the prevailing assumption that AI can substitute for traditional security auditing and provide critical evidence for defining the security boundaries of AI-assisted development and designing effective human–AI collaboration paradigms.
This study systematically investigates the security debt introduced by autonomous coding agents, which, while enhancing development efficiency, generate high-risk vulnerabilities often missed by conventional human code reviews. Leveraging the AIDev dataset, a validated LLM-as-a-judge evaluation framework, and qualitative human analysis, the authors find that 38.9% of agent-generated pull requests contain security smells, with 82.3% pertaining to supply chain integrity and 99.6% of critical smells involving hardcoded credentials. Notably, 81.1% of these credential leaks originate from human collaborators and evade detection by existing review mechanisms, exposing a critical blind spot in current development workflows.
研究通过扫描3,171个GitHub仓库,检测AI编码代理配置中的供应链缺陷,使用独立实现和语言模型仲裁验证结果。
本文研究了AI生成的Ansible代码是否满足安全要求,并提出了一种将Ansible最佳实践和CIS基准集成到提示中的方法,以预防安全问题。
This study addresses the high vulnerability rates and lack of real-world threat context in AI-generated code by proposing a just-in-time security remediation pipeline that integrates static analysis with large language models. Leveraging threat intelligence such as MITRE ATT&CK to enrich contextual understanding, the framework enables parallel vulnerability scanning, intelligent validation, and automated repair. Experimental results demonstrate that the pipeline reduces vulnerabilities by up to 69% with an 81% judgment consistency rate, while revealing a non-monotonic relationship between model capability and pipeline efficacy. These findings validate the effectiveness of knowledge-augmented automation in enhancing AI code security, with Sonnet 4.6 achieving optimal remediation performance.
This study addresses the security risks arising from AI-generated code that exceeds human review capacity and remains difficult for non-expert developers to govern. To this end, it proposes a multi-agent meta-agent system orchestrated by non-technical personnel. By integrating software agents, automated testing, and monitoring-auditing techniques, the system binds objectives, evidence, permissions, and decisions to a unified underlying goal, establishing a human-AI collaborative architecture for automated governance. The research demonstrates the unreliability of single-review mechanisms and reveals inherent limitations of monitoring agents. Furthermore, it establishes a paradigm of ultimate human control centered on objective alignment. The proposed framework is validated through deployment in a production-grade medical platform, achieving effective safety governance over AI-generated code in high-stakes environments.