Institution profile

AIM Intelligence

Research institution
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

Oct 05, 2026

This study addresses the limitation of traditional multimodal safety evaluations, which rely solely on behavioral classification and fail to reveal the controllability of models' internal safety mechanisms. To this end, we propose a "controllability profiling" framework that leverages sparse autoencoders (SAEs) to establish internal representation controllability as an independent evaluation dimension for the first time, quantifying both the detectability of safety signals and their intervention sensitivity in vision-language models. Implicit toxicity stress tests conducted on LlavaGuard and Qwen3.5 demonstrate significant discrepancies across models regarding the alignment between internal readout capabilities and selective control. By exposing these divergences, this work provides a critical theoretical foundation for developing next-generation multimodal safety benchmarks.

0 citationsRead paper

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

Oct 05, 2026

This study addresses the security-utility trade-off and architecture dependency inherent in existing decoding-stage defenses for large language models (LLMs) by proposing a cross-model universal defense method based on first-token dark knowledge. The approach mines latent safety signals within dark knowledge, constructing a model-agnostic defense direction through Top-k extraction and tokenizer mapping, while employing k-nearest neighbor (kNN) classification for jailbreak attack detection. Experimental results demonstrate that this framework significantly reduces attack success rates across diverse LLMs, effectively overcoming architecture-specific limitations and achieving an optimal balance between security and utility.

0 citationsRead paper

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Jun 18, 2026

This work addresses the insufficient robustness of large language model (LLM) agents in controlling safety-critical systems under persistent, adaptive adversarial attacks. To this end, the authors introduce NRT-Bench, a novel benchmark that simulates a nuclear power plant control room staffed by a five-member LLM operator team. The framework evaluates agent resilience through multi-channel, multi-turn red-teaming attacks coupled with an adversarial feedback mechanism. Crucially, it defines objective harm via the loss of critical safety functions grounded in actual system states—rather than textual judgments—and employs a fixed attack pairing replay protocol. Experiments across four state-of-the-art models reveal that 8.7%–12.1% of attack sessions result in safety function loss. While none of the 149 attacks compromised all models, approximately one-third succeeded against at least one, highlighting highly heterogeneous vulnerabilities and strong model-dependent defense efficacy.

0 citationsRead paper

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

May 26, 2026

This work addresses the “fragile safety” of language models, which mechanically adhere to original safety rules even when contextual shifts invert the safety implications of their actions. To systematically evaluate robustness in dynamic scenarios, we introduce a context-flipping assessment framework that constructs paired examples with reversed safety outcomes. Our analysis reveals, for the first time, a substantial gap—averaging 17.4 percentage points—between models’ safety reasoning and commonsense understanding, demonstrating that this fragility stems from insufficient policy coverage rather than misinterpretation. To mitigate this, we propose a state-aware verification mechanism that replaces conventional action-level safeguards. Evaluated on the PacifAIst benchmark and catastrophic consequence probes, our approach achieves 100% risk detection with zero false positives, whereas existing safeguards completely fail.

0 citationsRead paper

Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models

Mar 17, 2026

Current safety alignment mechanisms in diffusion-based language models assume that once a refusal token is generated, it remains immutable—a vulnerability this work exploits. We introduce TrajHijack, the first trajectory-level hijacking attack, which overwrites previously generated refusal tokens by re-masking them and injecting a fixed compliant prefix, enabling gradient-free, cross-model attacks. This approach exposes critical weaknesses in the dual-component safety architecture—comprising refusal detection and content generation—and reveals a counterintuitive defense inversion effect, wherein stronger defenses become more susceptible to attack. Evaluated on HarmBench, TrajHijack achieves attack success rates of 74–98%, with the state-of-the-art A2D defense exhibiting heightened vulnerability (89.9% ASR).

0 citationsRead paper
Recent publications

Latest Papers

Multimodal Safety Evaluation Should Measure Controllability Beyond Classification

Oct 05, 2026

This study addresses the limitation of traditional multimodal safety evaluations, which rely solely on behavioral classification and fail to reveal the controllability of models' internal safety mechanisms. To this end, we propose a "controllability profiling" framework that leverages sparse autoencoders (SAEs) to establish internal representation controllability as an independent evaluation dimension for the first time, quantifying both the detectability of safety signals and their intervention sensitivity in vision-language models. Implicit toxicity stress tests conducted on LlavaGuard and Qwen3.5 demonstrate significant discrepancies across models regarding the alignment between internal readout capabilities and selective control. By exposing these divergences, this work provides a critical theoretical foundation for developing next-generation multimodal safety benchmarks.

0 citationsRead paper

Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

Oct 05, 2026

This study addresses the security-utility trade-off and architecture dependency inherent in existing decoding-stage defenses for large language models (LLMs) by proposing a cross-model universal defense method based on first-token dark knowledge. The approach mines latent safety signals within dark knowledge, constructing a model-agnostic defense direction through Top-k extraction and tokenizer mapping, while employing k-nearest neighbor (kNN) classification for jailbreak attack detection. Experimental results demonstrate that this framework significantly reduces attack success rates across diverse LLMs, effectively overcoming architecture-specific limitations and achieving an optimal balance between security and utility.

0 citationsRead paper

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

Jun 18, 2026

This work addresses the insufficient robustness of large language model (LLM) agents in controlling safety-critical systems under persistent, adaptive adversarial attacks. To this end, the authors introduce NRT-Bench, a novel benchmark that simulates a nuclear power plant control room staffed by a five-member LLM operator team. The framework evaluates agent resilience through multi-channel, multi-turn red-teaming attacks coupled with an adversarial feedback mechanism. Crucially, it defines objective harm via the loss of critical safety functions grounded in actual system states—rather than textual judgments—and employs a fixed attack pairing replay protocol. Experiments across four state-of-the-art models reveal that 8.7%–12.1% of attack sessions result in safety function loss. While none of the 149 attacks compromised all models, approximately one-third succeeded against at least one, highlighting highly heterogeneous vulnerabilities and strong model-dependent defense efficacy.

0 citationsRead paper

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

May 26, 2026

This work addresses the “fragile safety” of language models, which mechanically adhere to original safety rules even when contextual shifts invert the safety implications of their actions. To systematically evaluate robustness in dynamic scenarios, we introduce a context-flipping assessment framework that constructs paired examples with reversed safety outcomes. Our analysis reveals, for the first time, a substantial gap—averaging 17.4 percentage points—between models’ safety reasoning and commonsense understanding, demonstrating that this fragility stems from insufficient policy coverage rather than misinterpretation. To mitigate this, we propose a state-aware verification mechanism that replaces conventional action-level safeguards. Evaluated on the PacifAIst benchmark and catastrophic consequence probes, our approach achieves 100% risk detection with zero false positives, whereas existing safeguards completely fail.

0 citationsRead paper

Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models

Mar 17, 2026

Current safety alignment mechanisms in diffusion-based language models assume that once a refusal token is generated, it remains immutable—a vulnerability this work exploits. We introduce TrajHijack, the first trajectory-level hijacking attack, which overwrites previously generated refusal tokens by re-masking them and injecting a fixed compliant prefix, enabling gradient-free, cross-model attacks. This approach exposes critical weaknesses in the dual-component safety architecture—comprising refusal detection and content generation—and reveals a counterintuitive defense inversion effect, wherein stronger defenses become more susceptible to attack. Evaluated on HarmBench, TrajHijack achieves attack success rates of 74–98%, with the state-of-the-art A2D defense exhibiting heightened vulnerability (89.9% ASR).

0 citationsRead paper