Institution profile

Qihoo 360

Industry researchasia · cn
Official website
Research library39linked papers
Opportunities0open roles
Selected work

Representative Papers

From Latents to Wires: Surgical Post-Editing on Large Language Models

Sep 26, 2026

This study addresses the challenge of precisely localizing and removing specific semantic targets—such as watermarks and identity claims—from large language models (LLMs) while preserving their general capabilities. To this end, it proposes L2W, a novel post-training "post-editing" paradigm that transcends the limitations of conventional fine-tuning. The framework employs Jacobian lens attribution analysis to identify critical components, integrating counterexample-guided causal ablation with a cumulative deactivation strategy to achieve surgical-precision editing. Empirically, this approach successfully eliminates implanted watermarks, metadata self-claims, and adult-content refusal behaviors. Furthermore, it demonstrates composite dual editing in text-to-image models, validating the effectiveness of lossless, precise removal.

0 citationsRead paper

SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills

Aug 04, 2026

This work addresses the challenge that external skills invoked by large language model (LLM) agents may exhibit latent harmful behaviors under specific environmental conditions or interaction histories—risks that evade detection by existing static analysis methods. To tackle this, we propose SkillSentry, a dynamic security testing framework that simulates bait environments using LLMs, adaptively generates exploratory tasks, and compares execution trajectories with and without the target skill enabled. By correlating source code and runtime logs, SkillSentry enables precise attribution and detection of conditionally triggered malicious behaviors, overcoming the limitations of static approaches. Empirical evaluation shows that SkillSentry achieves 99.50% recall and an average F1 score of 96.26% on standard benchmarks. Notably, under semantic-preserving evasion attacks, it maintains robust performance with an average F1 of 92.95%, substantially outperforming the strongest baseline at 80.07%.

0 citationsRead paper

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

Jul 20, 2026

Current jailbreaking attacks on text-to-image models suffer from low efficiency, semantic collapse, and neglect of critical information in defense feedback. This work proposes the MIND framework, which, for the first time, formulates jailbreaking as a cognitive reasoning process over latent defense mechanisms. By integrating multimodal feedback parsing, dynamically updating a defense profile, and employing meta-memory-driven strategy retrieval, MIND enables semantically coherent and adaptive attacks. The approach transcends the limitations of conventional black-box optimization, achieving a 95.62% attack success rate (ASR) across six defense configurations on Stable Diffusion v1.5 and up to 91.58% ASR on four major commercial text-to-image generation systems.

0 citationsRead paper

ElephantAgent: Contextual State Continuity in Agentic Systems

Jul 02, 2026

This work addresses the vulnerability of intelligent agent systems to context-state poisoning attacks—stemming from their reliance on external tools and memory—and the absence of verifiable guarantees for state continuity. To mitigate these risks, the authors propose ElephantAgent, a novel protocol that introduces state continuity mechanisms into dynamic context management for agent systems. By recomputing and verifying a digest of the local context state prior to each query and leveraging trusted hardware to maintain a linearized log of authorized state transitions, ElephantAgent ensures historical traceability. This approach effectively defends against attacks such as tool descriptor tampering and memory poisoning, while enabling anomaly detection and rollback to known-good states.

0 citationsRead paper
Recent publications

Latest Papers

From Latents to Wires: Surgical Post-Editing on Large Language Models

Sep 26, 2026

This study addresses the challenge of precisely localizing and removing specific semantic targets—such as watermarks and identity claims—from large language models (LLMs) while preserving their general capabilities. To this end, it proposes L2W, a novel post-training "post-editing" paradigm that transcends the limitations of conventional fine-tuning. The framework employs Jacobian lens attribution analysis to identify critical components, integrating counterexample-guided causal ablation with a cumulative deactivation strategy to achieve surgical-precision editing. Empirically, this approach successfully eliminates implanted watermarks, metadata self-claims, and adult-content refusal behaviors. Furthermore, it demonstrates composite dual editing in text-to-image models, validating the effectiveness of lossless, precise removal.

0 citationsRead paper

SkillSentry: Adaptive Honey Worlds for Dynamic Safety Testing of Agent Skills

Aug 04, 2026

This work addresses the challenge that external skills invoked by large language model (LLM) agents may exhibit latent harmful behaviors under specific environmental conditions or interaction histories—risks that evade detection by existing static analysis methods. To tackle this, we propose SkillSentry, a dynamic security testing framework that simulates bait environments using LLMs, adaptively generates exploratory tasks, and compares execution trajectories with and without the target skill enabled. By correlating source code and runtime logs, SkillSentry enables precise attribution and detection of conditionally triggered malicious behaviors, overcoming the limitations of static approaches. Empirical evaluation shows that SkillSentry achieves 99.50% recall and an average F1 score of 96.26% on standard benchmarks. Notably, under semantic-preserving evasion attacks, it maintains robust performance with an average F1 of 92.95%, substantially outperforming the strongest baseline at 80.07%.

0 citationsRead paper

Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models

Jul 20, 2026

Current jailbreaking attacks on text-to-image models suffer from low efficiency, semantic collapse, and neglect of critical information in defense feedback. This work proposes the MIND framework, which, for the first time, formulates jailbreaking as a cognitive reasoning process over latent defense mechanisms. By integrating multimodal feedback parsing, dynamically updating a defense profile, and employing meta-memory-driven strategy retrieval, MIND enables semantically coherent and adaptive attacks. The approach transcends the limitations of conventional black-box optimization, achieving a 95.62% attack success rate (ASR) across six defense configurations on Stable Diffusion v1.5 and up to 91.58% ASR on four major commercial text-to-image generation systems.

0 citationsRead paper

ElephantAgent: Contextual State Continuity in Agentic Systems

Jul 02, 2026

This work addresses the vulnerability of intelligent agent systems to context-state poisoning attacks—stemming from their reliance on external tools and memory—and the absence of verifiable guarantees for state continuity. To mitigate these risks, the authors propose ElephantAgent, a novel protocol that introduces state continuity mechanisms into dynamic context management for agent systems. By recomputing and verifying a digest of the local context state prior to each query and leveraging trusted hardware to maintain a linearized log of authorized state transitions, ElephantAgent ensures historical traceability. This approach effectively defends against attacks such as tool descriptor tampering and memory poisoning, while enabling anomaly detection and rollback to known-good states.

0 citationsRead paper