SWE-PolyVision: Benchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering

πŸ“… 2026-09-24
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the lack of benchmarks evaluating agents’ ability to integrate multi-image evidence for repository-level code repair. To this end, we construct an execution-based benchmark comprising 92 real-world tasks to systematically assess models’ cross-image reasoning and repair capabilities under multimodal conditions. Methodologically, we decouple the availability of multi-image evidence from its effective utilization, avoiding the conflation of patch success with explicit reasoning. Experiments are conducted across static images, videos, and text-only inputs, employing native visual and tool-mediated access modalities. Results indicate that the impact of visual access on task performance varies significantly across models and tasks, and current multimodal translation mechanisms remain unstable.
πŸ“ Abstract
Current multimodal software-engineering benchmarks expose images as additional context, but do not test whether an agent can integrate evidence distributed across images into a verified repository-level repair. We present SWE-PolyVision, an executable benchmark of 92 real tasks from 36 open-source organizations, with 48 public tasks and 44 private holdouts. The release contains 402 static images and 6 videos, with at least two visual inputs per task. Each task pairs a fixed pre-fix repository with an isolated verifier and is evaluated under the supported conditions among three access modes: Text-only, Native Vision, and Tool-mediated Vision. Across eleven coding models, visual access changes which tasks are solved, but effects depend on both model and task. Two trace-linked Native Vision cases illustrate how complementary visual and textual clues can lead to source-localized, verified repairs; controlled interventions show that this conversion is not yet stable across inputs. SWE-PolyVision thus separates the availability of multi-image evidence from its successful use in repository-level repair, without treating patch success alone as proof of explicit reasoning.
Problem

Research questions and friction points this paper is trying to address.

cross-image abductive reasoning
repository-level software engineering
multimodal benchmark
visual evidence integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Image Abductive Reasoning
Multimodal Benchmark
Repository-Level Software Engineering
Tool-mediated Vision
Executable Evaluation