VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation in existing multimodal generation benchmarks regarding the joint capability of handling multiple reference images and heterogeneous visual instructions. We propose the first joint evaluation framework for multi-reference and heterogeneous visual instructions, constructing a benchmark comprising 1,241 tasks that encompass competitive scenarios and image-text comparisons. Through controlled variable experiments, we systematically assess model generation performance under complex conditions. Our findings reveal a trade-off mechanism between instruction following and artifact generation: strong instruction adherence exacerbates artifacts, salient state references diminish following fidelity, and direct visual constraints outperform textual descriptions. This work bridges the gap in multi-condition joint evaluation, providing critical empirical evidence for controllable image generation.
📝 Abstract
Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7) under multiple heterogeneous visual instructions (up to 6), (ii) cases where reference images can potentially compete with visual instructions (e.g., a strongly posed subject vs. a target pose), and (iii) controlled comparison of visual instructions with text descriptions at different levels of specificity. Using these capabilities, we uncover three findings: (1) models face an adherence-artifact trade-off: once models reach stronger visual instruction adherence, stronger adherence tends to coincide with more instruction artifacts in generated images, (2) visual instruction adherence tends to be lower on tasks whose reference images carry a salient state of the controlled attribute (e.g., a neon-lit subject under a light-direction instruction), most consistently for light and wind, and (3) for models that can understand visual instructions, it is often better to provide visual constraints directly rather than describe them in text; when using text, a moderate level of detail works better than an exhaustive description. VIF-Bench is released as an open benchmark to establish a basis for fair comparison in controllable multi-reference image generation.
Problem

Research questions and friction points this paper is trying to address.

multi-reference image generation
visual instruction following
benchmark evaluation
controllable generation
multimodal generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Instruction Following
Multi-Reference Image Generation
Benchmark Evaluation
Controllable Generation
Multimodal Models
🔎 Similar Papers
No similar papers found.