henry 发自 凹非寺
Qbit | Public Account QbitAI
The visual native Harness for multi-modal models is here!
Recently, He Kai Ming’s team presented a visual-native Harness in their latest paper titled “VISTA: A Visual Harness for Reasoning in an Interactive World”.
VISTA.
It allows existing multimodal models to directly observe the environment, preserve the original images, and review, zoom in, and check details at any time during reasoning, without having to be trained from scratch.
Compared to the traditional approach of abstracting the environment into text or code, VISTA preserves the original visual information. The model can, based on current issues, re-insert past images into the context and participate in current judgment and reasoning.
Based on this approach, visual experience becomes a context that the model can call at any time and examine repeatedly, and multi-modal models also have a native scaffold built around visual memory.
In terms of evaluation, with VISTA and Claude Opus 5.0, all 25 public games of ARC-AGI-3 were completed, achieving a full score of 100 in relative human action efficiency. The number of game actions used was 57.4% less than the human baseline for the first play.
Switched to GPT-5.6 Sol, also completed all levels, achieving a score of 99.
Furthermore, the team has also tested this approach in browser games, mazes, and connection-based problems, and has identified embodied tasks that are closer to the physical world as future research directions.
How is this done?
Visual-native scaffolding
From ResNet, Mask R-CNN to MAE, visual perception and representation learning have always been a main thread in He Kai-ming’s research.
This time, VISTA has turned its attention to a new problem: how to enable the model to accumulate, store, and utilize visual experience during continuous interaction?
The answer is Harness.
In November last year, Anthropic used a long-duration programming task as an example to show how Harness helps models cross multiple context windows and continue advancing the task.
Specifically, Harness will record tasks that have not been completed in a task list, and save the completed work in progress files and code.
In this way, even when the model enters a new context window, it can understand what has been done before and what should be done next by reading these information.
Complex tasks can be executed step by step, checked for errors, and continue working between different context windows.
It can be said that this Harness, which saves task status primarily through text and code, has promoted the emergence of Agent capabilities to a certain extent.
However, in a real-world environment, just having verbal memory is obviously not enough.
Because once entering a real visual environment, the model not only needs to remember the position and orientation of objects, but also needs to understand what changes occur in the environment before and after each action.
Moreover, when a model sees a picture for the first time, it may not know which details will become important later on.
At the same time, it is also difficult to fully record these visual information in a written note in advance, and it is even more challenging to ensure that the recorded information remains sufficient during the model’s subsequent tasks.
In benchmarks like ARC-AGI-3, one approach is to first convert the image into a text grid composed of numbers, and then have the model write programs to simulate environment rules and verify action plans.
(Note: ARC-AGI-3 is a interactive visual reasoning benchmark that does not provide pre-defined rules or goals; the agent needs to observe and act to figure out how to complete the task on its own.)
But the problem is that the more complex the environment, the harder it is to fully convert the appearance of objects, spatial relationships, and dynamic changes into text and code. Moreover, as the interaction history grows longer, early images will gradually be removed from the context, or replaced by text summaries.
So, here, the research question of VISTA becomes:
How can an agent re-examine visual information it has seen in the past when reasoning is needed, without additional training of a base model?
Based on this, VISTA stores each frame returned by the environment exactly as it is, outside the context, including the intermediate frames during the animation process. The model retrieves them again when needed.
In the ARC-AGI-3 evaluation, it can compare images at different times, zoom in on specific areas, and check the orientation markers on small squares.
Among them, the text note recording model records the current understanding, while the original image retains the evidence for it to be re-evaluated.
The research refers to this mechanism as explicit attention to the interaction history. The model selects which visual information should re-enter the context based on the questions being considered.
To make this choice possible, the system must save the original image and provide tools that can be accessed and examined at any time.
Specific implementation
To enable the model to save and utilize past visual experiences, VISTA has designed three core components:
Visual observation, lossless visual memory, and active visual inspection. These three components have distinct roles: one is responsible for making the model see the image clearly, another for preserving history, and the third for reviewing as needed.
First is visual observation, which is responsible for providing the environmental image to the model.
In the ARC-AGI-3 experiment, VISTA will scale the official 64×64 image to a 512×512 PNG image, while preserving the object’s colors, appearance, and spatial relationships, allowing the model to directly observe the environment.
Next is the lossless visual memory, which is responsible for storing the images seen by the model.
After each action is executed, VISTA will completely save all the images returned by the environment, including the intermediate frames of the animation during the action process, and establish an index based on the round number and frame number.
In this way, the model does not need to decide in advance which information is worth remembering, but can instead fully save the visual experience for later use.
Finally, the active visual check is responsible for making the model review historical frames as needed.
Through the inspect tool, the model can specify a certain historical round, retrieve the corresponding image, or crop and zoom in on specific areas to examine the details.
If you want to know exactly what a certain operation changed, the model can also retrieve multiple frames at a time to compare the changes before and after the action.
The entire process is equivalent to providing the model with a visual archive that can be reviewed at any time. It not only sees the current environment, but also actively looks back at past observations, providing a basis for future actions.
Speaking of this, how do these three components work together specifically?
According to the VISTA execution process, at the start of each round, the model first observes the current scene and available actions, then combines the experience accumulated before to determine the state of the environment, and proposes assumptions for the next action.
If the available information is insufficient, the model can call tools to review historical images, zoom in on specific areas, and even retrieve detailed pixel information in order to find more evidence.
After finding evidence, the model does not act directly, but first predicts what changes might result from this operation.
After all the actions are completed, compare the actual results with the predictions, check whether one’s judgment is correct, and accordingly adjust the understanding of the game rules.
Repeatedly, the model can explore the environment while accumulating experience during continuous interaction.
Of course, just visual files are not enough. To ensure the model remains coherent during long-term tasks, VISTA also introduced two written notes:
- GUIDE.md: Record the rules and experiences that can be reused between different levels.
- WORKING.md: Records the status, progress, and next steps of the current level.
When the context approaches the upper limit, the model first organizes the transition summary, and then enters a new context window to continue executing tasks. The previous text notes, action history, and visual files will all be retained.
There is another noteworthy design: although VISTA preserves all historical images intact, it does not cram all the pictures into the context of the model.
In the complete solution, after each action is executed, the framework defaults to displaying only the last frame for the model. As for the intermediate images during the action, they are uniformly saved in a visual file, and can be actively retrieved by the model when needed.
This way, the complete visual information is preserved, and a large number of historical images from the context do not occupy space.
More importantly, VISTA did not train a new model additionally for this purpose.
Environmental understanding, action planning, and reasoning are still performed by ready-made multimodal models, while Harness is responsible for executing tool calls, saving visual history, and providing relevant evidence when the model needs it.
In other words, what VISTA changes is not the model itself, but the way the model acquires, stores, and uses visual information.
Experimental validation
In addition to the experiments mentioned at the beginning, the team also tested VISTA on three benchmarks: GameWorld, AI GameStore, and BabyVision. It covered tasks such as browser games, mazes, and connections.
Using the same GPT-5.6 Sol model, compared to the official base framework, VISTA increased the success rate from 40.0% to 63.3% in 170 tasks of GameWorld.
Of the 10 games on AI GameStore, the overall score increased from 47.3 to 140.3, where the human median is normalized to 100;
On the 39 mazes and connection puzzles selected by BabyVision, the accuracy increased from 41.0% to 63.2%. Finally, this set of tasks consisted only of static images, and the model could improve its judgment by magnifying specific parts and checking pixels.
These results indicate that preserving the original image and allowing the model to view it as needed can further leverage the capabilities of existing multimodal models.
The same VISTA framework only requires minimal adaptation and can be used for different games and puzzles, showing its potential in more visual tasks.
In conclusion, the research points toward more physical-world-aligned tasks for future validation: enabling models to store, review, and utilize their visual experiences in a continuously changing environment, and decide on the next action.
Author Introduction
Finally, let us introduce the authors of this work.
The three co-first authors of the paper are Qiushi Han, Hu Keya, and Linlu Qiu. The other two authors are Cathy Wu and He Kai Ming.
Qiushi Han (Josh Han) is currently a doctoral student at the MIT Operations Research Center, under the guidance of Cathy Wu.
Keya Hu is currently a doctoral student in Electrical Engineering and Computer Science at MIT, guided by He Kaiming and Jacob Andreas.
She graduated from the ACM class at Shanghai Jiao Tong University with a bachelor’s degree. Her research interests focus on the intersection of language and vision, and she hopes to develop an AI agent that is more efficient in terms of data usage and has stronger generalization capabilities.
Linlu Qiu is currently a doctoral student in the Department of Electrical Engineering and Computer Science at MIT, as well as in the Laboratory of Computer Science and Artificial Intelligence. His advisors are Yoon Kim and Jacob Andreas.
Its research areas include natural language processing and machine learning, and it has conducted research at Google Research and Meta FAIR.
Cathy Wu is currently an associate professor in the Department of Civil and Environmental Engineering and the Institute for Data Systems and Society at MIT. She studies how machine learning and reinforcement learning can improve complex systems such as transportation.
She received a bachelor’s and engineering master’s degree from MIT, and then a doctorate from the University of California, Berkeley.
He Kai-ming is currently a tenured associate professor in the Department of Electrical Engineering and Computer Science at MIT, and he is also the principal author of works such as ResNet, Mask R-CNN, and MAE.
He graduated from Tsinghua University with a bachelor’s degree and from The Chinese University of Hong Kong with a doctorate. He worked at Microsoft Research Asia and FAIR, and joined MIT in 2024.
Reference link
[1] https://arxiv.org/pdf/2610.02200