Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents
本文针对AI代理安全事件报告问题,通过专家意见识别必要信息,并提出高效记录事件和评估漏洞泛化等研究方向。
本文针对AI代理安全事件报告问题,通过专家意见识别必要信息,并提出高效记录事件和评估漏洞泛化等研究方向。
This work systematically uncovers a critical compatibility vulnerability in encrypted reasoning chains employed by large language model (LLM) providers to safeguard intellectual property. Despite encryption, reasoning blocks exhibit cross-session, cross-user, and cross-model interoperability flaws. The authors propose a novel decryption-based jailbreaking technique that leverages reverse engineering and cross-model reasoning block injection: by exploiting weaker models to decrypt the encrypted inference traces of stronger ones, the method reconstructs full reasoning processes without direct attacks. Empirical evaluation demonstrates successful extraction of reasoning chains from models by Anthropic, OpenAI, and Google, decrypting 315,320 publicly logged blocks, recovering 367 personally identifiable information (PII) instances and 182 credential sets, and enabling stealthy prompt injection that effectively bypasses existing anti-distillation and security mechanisms.
This study systematically evaluates the reproducibility and stability of large language models (LLMs) in repeatedly performing JavaScript security scanning tasks. Drawing on 300 repeated experiments using LLMs such as Claude, the Snyk Code static analysis tool, and a newly developed benchmark—Snyk VulnBench JS 1.0—the work provides the first quantitative evidence that LLMs exhibit high stability in detecting reference vulnerabilities (134 out of 158 consistently reproduced within five trials), while showing significant inconsistency for non-matching vulnerabilities (only 22 out of 161 stably detected). Beyond uncovering potential blind spots in Snyk Code, the research proposes a novel paradigm of complementary collaboration between LLMs and static application security testing (SAST), demonstrating that hybrid approaches effectively enhance the comprehensiveness of vulnerability detection.
This study addresses critical security risks in mainstream AI agent skill marketplaces, where malicious payloads and severe vulnerabilities endanger user credentials and system integrity. Conducting the first large-scale security audit of 3,984 skills across three dominant platforms, the authors combine automated static analysis with manual dynamic validation to establish the first threat taxonomy and attack pattern framework tailored to the AI skill ecosystem. Their investigation identifies 76 malicious payloads and reveals that 13.4% of examined skills contain high-severity vulnerabilities. Notably, at least eight malicious skills remained publicly accessible at the time of publication, underscoring significant security gaps and insufficient oversight within the current AI agent marketplace landscape.
本文针对AI代理安全事件报告问题,通过专家意见识别必要信息,并提出高效记录事件和评估漏洞泛化等研究方向。
This work systematically uncovers a critical compatibility vulnerability in encrypted reasoning chains employed by large language model (LLM) providers to safeguard intellectual property. Despite encryption, reasoning blocks exhibit cross-session, cross-user, and cross-model interoperability flaws. The authors propose a novel decryption-based jailbreaking technique that leverages reverse engineering and cross-model reasoning block injection: by exploiting weaker models to decrypt the encrypted inference traces of stronger ones, the method reconstructs full reasoning processes without direct attacks. Empirical evaluation demonstrates successful extraction of reasoning chains from models by Anthropic, OpenAI, and Google, decrypting 315,320 publicly logged blocks, recovering 367 personally identifiable information (PII) instances and 182 credential sets, and enabling stealthy prompt injection that effectively bypasses existing anti-distillation and security mechanisms.
This study systematically evaluates the reproducibility and stability of large language models (LLMs) in repeatedly performing JavaScript security scanning tasks. Drawing on 300 repeated experiments using LLMs such as Claude, the Snyk Code static analysis tool, and a newly developed benchmark—Snyk VulnBench JS 1.0—the work provides the first quantitative evidence that LLMs exhibit high stability in detecting reference vulnerabilities (134 out of 158 consistently reproduced within five trials), while showing significant inconsistency for non-matching vulnerabilities (only 22 out of 161 stably detected). Beyond uncovering potential blind spots in Snyk Code, the research proposes a novel paradigm of complementary collaboration between LLMs and static application security testing (SAST), demonstrating that hybrid approaches effectively enhance the comprehensiveness of vulnerability detection.
This study addresses critical security risks in mainstream AI agent skill marketplaces, where malicious payloads and severe vulnerabilities endanger user credentials and system integrity. Conducting the first large-scale security audit of 3,984 skills across three dominant platforms, the authors combine automated static analysis with manual dynamic validation to establish the first threat taxonomy and attack pattern framework tailored to the AI skill ecosystem. Their investigation identifies 76 malicious payloads and reveals that 13.4% of examined skills contain high-severity vulnerabilities. Notably, at least eight malicious skills remained publicly accessible at the time of publication, underscoring significant security gaps and insufficient oversight within the current AI agent marketplace landscape.