VISTA: A Visual Harness for Reasoning in an Interactive World

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottleneck of long-horizon visual memory loss faced by multimodal models in interactive environments. To this end, we propose VISTA, a universal visual framework that incorporates a lossless visual memory mechanism and an active observation retrieval technique. By preserving raw observations and dynamically reorganizing inputs, VISTA transcends the short-term memory limitations of conventional agents, endowing models with complex interactive reasoning capabilities across diverse environments. Experimental results demonstrate that VISTA enables Claude Opus 5.0 to achieve perfect performance on the ARC-AGI-3 benchmark while reducing action steps by 57.4%, significantly outperforming existing baseline methods.
📝 Abstract
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
Problem

Research questions and friction points this paper is trying to address.

multimodal reasoning
interactive environments
long-horizon vision
visual memory
multimodal agents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Harness
Multimodal Reasoning
Lossless Visual Memory
Interactive Environments
Long-horizon Vision
🔎 Similar Papers
No similar papers found.