A$^2$Safe: Counterfactual Evidence-Aligned Adaptive Agent Collaboration for Safe and Effective Visual Question Answering

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉问答中基于表面关联而非实际证据的安全决策问题,提出A²Safe框架,通过反事实证据对齐和自适应协作提高安全性和有效性。
📝 Abstract
Visual Question Answering (VQA) with Multimodal Large Language Models (MLLMs) requires not only producing safe and effective responses, but also grounding safety decisions in the multimodal evidence that determines risk. Recent safety-alignment methods improve refusal behavior and contextual risk awareness, yet correct safety outcomes may still rely on superficial textual or visual correlations, particularly when risk emerges from interactions between individually benign image and question content. To address this issue, we propose A$^2$Safe, a counterfactual evidence-aligned adaptive agent collaboration framework for safe and effective VQA. A$^2$Safe organizes localized visual observations, textual intent, and cross-modal risk relations through a Grounded Safety Evidence Board, making the basis of safety decisions explicit. Counterfactual safety evidence alignment enforces invariance to safety-irrelevant changes while requiring appropriate safety-state and response-mode transitions when risk-critical evidence is minimally altered. The resulting evidence state further supports adaptive collaboration, enabling direct answering when grounded evidence is sufficient and invoking policy critique and response revision when evidence is risky, uncertain, or conflicting. Under complementary safety-critical and general VQA protocols, A$^2$Safe achieves a 95.72 SIUO safety score, reduces the benign refusal rate on MOSSBench to 14.67%, and maintains an average general VQA score of 78.34 with 27.8% token overhead. These results support counterfactual evidence-aligned adaptive collaboration for safe and effective multimodal question answering.
Problem

Research questions and friction points this paper is trying to address.

Visual Question Answering
Multimodal Large Language Models
Safety Alignment
Counterfactual Evidence
Adaptive Collaboration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Evidence-Aligned
Adaptive Agent Collaboration
Visual Question Answering
Safety Alignment
Multimodal Large Language Models
💼 Related Jobs
No related jobs found.
Q
Quanxing Xu
School of Computer Science and Engineering, Macau University of Science and Technology, Macao SAR 999078, China
L
Ling Zhou
School of Computer Science and Engineering, Macau University of Science and Technology, Macao SAR 999078, China
X
Xian Zhong
Hubei Key Laboratory of Transportation Internet of Things, School of Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China
Jinyu Tian
Jinyu Tian
Macau University of Science and Technology
Adversarial Machine Learning
Xiaohua Huang
Xiaohua Huang
The University of Memphis
Cancer Nanomedicine
Rubing Huang
Rubing Huang
Macau University of Science and Technology
AI for Software EngineeringSoftware Engineering for AISoftware TestingAI Applications
C
Chia-Wen Lin
Department of Electrical Engineering, National Tsing Hua University, Hsinchu 30013, Taiwan