Visual Abstention in Unified Multimodal Models

๐Ÿ“… 2026-10-06
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the inability of unified multimodal models to recognize and refuse infeasible visual editing requests. It formally defines the concept of โ€œvisual abstentionโ€ and reveals a decoupling between editing capability and refusal behavior. To bridge this gap, the work proposes the DoD evaluation benchmark and the VisTA paired training strategy, which leverages supervised fine-tuning on both feasible and infeasible samples to enable autonomous decision-making without explicit prompting. Furthermore, it establishes a joint evaluation framework that simultaneously measures editing success rate and refusal accuracy. The resulting VisTA-BAGEL model achieves a 93.0% refusal rate and a 74.3% editing completion rate, significantly outperforming existing mainstream models and demonstrating an effective balance between precise refusal and efficient editing.
๐Ÿ“ Abstract
Unified multimodal models (UMMs) integrate understanding and generation, yet their generative behavior is rarely governed by what they understand about the task. We formalize visual abstention: when a requested visual transformation is impossible under the task's rules, the model should recognize that no valid solution exists, state this, and decline to generate. We introduce Draw-or-Decline (DoD), a benchmark of 1,050 feasible-infeasible request pairs across 7 task categories that jointly measures editing success and the refusal of infeasible requests. Evaluating 8 UMMs, we find that editing ability and abstention are distinct capabilities: even the strongest editor, at 68.4% editing accuracy, refuses only 0.4% of infeasible requests under ordinary instructions. Their reasoning shows why: the models rarely notice the conflict, and instead plan the edit as if the request were possible, often describing objects that are not in the image, or quietly change the request into one they can complete. Explicitly prompting these UMMs to report infeasibility increases textual refusals but reduces editing accuracy. We propose VisTA (Visual Transformation and Abstention), a training method that pairs feasible and infeasible examples so that a model judges feasibility before deciding whether to generate. We train VisTA-BAGEL to perform feasible edits and decline infeasible requests. Without any reminder, it refuses 93.0% of infeasible requests, up from 0.4% for the strongest editor, while falsely refusing only 0.8% of feasible ones. Unlike a reminder, this does not cost editing accuracy: VisTA-BAGEL completes 74.3% of feasible edits, more than any of the 8 evaluated UMMs.
Problem

Research questions and friction points this paper is trying to address.

Visual Abstention
Unified Multimodal Models
Image Editing
Infeasible Requests
Refusal Behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Abstention
Unified Multimodal Models
VisTA
Draw-or-Decline Benchmark
Feasibility Judgment
๐Ÿ”Ž Similar Papers
No similar papers found.