We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
Visual Perception and Lossless Memory
VISTA gives a general-purpose multimodal model long-horizon vision by allowing direct perception of the environment through visual observations. The harness maintains a lossless visual memory that preserves past observations in their original form. This design ensures that no detail is discarded during reasoning, providing a complete visual history that the model can reference at any point during task execution.
Active Retrieval and Reorganization
The model can actively retrieve observations from visual memory and reorganize its visual input as it reasons. This dynamic interaction between perception and recall allows the model to focus on relevant prior states without retracing steps or relying on compressed abstractions that may lose critical information.
ARC-AGI-3 Results
On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00. The model completed all 25 public games using 57.4% fewer actions than first-time human participants. This represents a substantial gain in both efficiency and completeness compared to the baseline model without the harness.
Cross-Benchmark Generalization
VISTA's simple design allows natural extension to diverse visual environments with minimal adaptation. Across three additional benchmarks covering visual games and puzzles, VISTA substantially outperforms baselines using the same underlying model with minimal harnesses. These results suggest the harness is not ARC-AGI-3-specific but offers broad improvements across interactive visual tasks.
Design Principles
The harness design prioritizes minimal complexity while maximizing performance gains. By focusing on lossless visual memory and active retrieval, VISTA avoids the computational overhead of more complex architectures while still unlocking significant reasoning capabilities in multimodal models.