ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of detecting implicit risks in multimodal large language models and mitigating unimodal shortcut learning. We introduce TriggerBench, the first dataset designed to model risk compositionality, and propose a step-supervised structured reasoning framework alongside a dedicated safeguard model, ThinkingGuard. Our approach decouples risk identification into progressive cognitive stages to eliminate residual risks. By integrating counterfactual contrast, step-reward Monte Carlo Tree Search, dual-constrained preference alignment, and knowledge distillation, the method enforces rigorous cross-modal logical deduction. Experimental results demonstrate that the proposed framework achieves superior performance on both standard and implicit safety benchmarks, substantially enhancing model safety and reliability.
📝 Abstract
While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the first dataset explicitly modeling risk compositionality (5,600 instances). By formally isolating Key Elements and Trigger Elements to build counterfactual contrastive pairs, TriggerBench eliminates risk residues and forces models to perform genuine logical deduction rather than superficial pattern matching, which provides a rigorous foundation for both large-scale training and fine-grained evaluation. Building on this, we propose a Step-Supervised Structured Reasoning training framework and employ it to train ThinkingGuard, a specialized guard model. Inspired by Situation Awareness theory, we decouple implicit risk identification into progressive cognitive stages, and utilize a step-reward Monte Carlo Tree Search algorithm to explore optimal reasoning trajectories, which are then distilled into the model through Dual-Constraint Preference Alignment. Extensive experiments across both standard and implicit safety benchmarks demonstrate that ThinkingGuard achieves strong performance. Project resources are available at https://github.com/FroggyChen/ThinkingGuard.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Implicit Risks
Safety Detection
Cross-modal Risk Activation
Shortcut Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Implicit Risks
Step-Supervised Structured Reasoning
Monte Carlo Tree Search
Counterfactual Contrastive Pairs
Dual-Constraint Preference Alignment
🔎 Similar Papers
No similar papers found.