量子位

New Work from Kaiming He's Team: Learning the ARC Challenge Just by Watching Cat Videos

Who would have thought that the one teaching AI to solve ARC abstract reasoning puzzles would be a cat? He Kaiming's team's latest paper proposesNAT-ARC, a purely visual ARC solving solution. It doe

Image source · 量子位

Who would have thought that the one teaching AI to solve ARC abstract reasoning puzzles would be a cat?

He Kaiming's team's latest paper proposesNAT-ARC, a purely visual ARC solving solution.

It doesn't rely on LLMs; instead it uses pretraining on cats, dogs, flowers and plants from ImageNet to teach the model to "see", and then transfers that to abstract grid reasoning.

As a result, NAT-ARC's best single model achieved a pass@2 score of 63.4% on ARC-1, rising to 70.2% after ensembling.

This is the first time a pure vision approachGetting closer to dedicated LLM systemsthe level of.

ARC, widely recognized as one of the hardest AI visual intelligence tests, gives you several input-output examples of colored grids, asks you to infer the hidden transformation rules behind them, and then apply them to new grids.

The mainstream approaches of the past all uniformly translated grids into text or symbols and handed them to an LLM for reasoning.

This paper goes against the grain — it doesn't even use ARC grids during the pretraining stage.

The visual abilities the model learned from observing the real world have transferred, quite effortlessly, to reasoning about entirely abstract colored grids.

The ARC track: vision-based approaches rise

ARC, short for Abstraction and Reasoning Corpus, was proposed in 2019 by François Chollet, the creator of Keras.

Every ARC puzzle has a uniform format: a 2D grid of at most 30×30 cells with up to 10 colors. Each puzzle provides 2 to 5 pairs of input-output examples, and the challenger must infer the hidden transformation rule from the examples, then apply that rule to a brand-new input grid to predict its output.

These puzzles sound like children's brain teasers about finding patterns in pictures, but ARC is extremely difficult for AI.

For one thing, every puzzle has a different transformation rule, so there is no fixed template to memorize.

There is also very little data: only two or three examples are given to induce an abstract rule and execute it precisely.

This combination has made ARC abenchmark for measuring AI abstract reasoning ability。

Currently, nearly all the dominant approaches on the ARC track are large language model solutions, whose common idea is to translate the grid into text or symbolic sequences and hand them to a language model.

The vision route only began to rise later.

A group of researchers took a completely different path: instead of translating grids into text, they used vision models to process grid images directly.

The most important step among these came from the same MIT team, which proposedVARC. This method redefines ARC as a conditional image-to-image translation task, achieving 54% with a 19M-parameter vision model.

LoopViT added a recurrent reasoning mechanism on top of this framework, reaching 65.8% with 18M parameters.

Loop-OWM used a video pretrained model for few-shot learning, reaching 68.5% with 10.6M parameters.

These models are several orders of magnitude smaller in parameter count than the LLM approaches, but their results are starting to catch up.

Yet the vision route still has a structural shortcoming.

One of the core reasons LLMs keep improving on ARC is that, through pretraining on massive corpora, they accumulate general capabilities — the larger the model and the more thorough the pretraining, the more evident the scaling on downstream tasks.

However, most visual ARC approaches train their models from random initialization and never enjoy the benefits of pretraining.

On this basis, the goal of NAT-ARC is to add the pretraining step to the visual route, attempting to break through its scaling bottleneck.

Done with cats, on to the ARC challenge

The core idea of NAT-ARC is toinsert an ImageNet MAE pretraining step in front of VARC's visual pipeline.。

MAE, Masked Autoencoder, is a self-supervised visual pretraining method proposed by Kaiming He during his time at FAIR in 2022.

Its training process works by masking most of a picture's regions and having the model reconstruct the original image from the remaining fragments.

Through this task, the model can understand the shape, structure, and spatial relationships of objects.

NAT-ARC's approach is to directly take the encoder weights that MAE was trained on ImageNet (a dataset containing about 1.3 million natural images of cats, dogs, flowers, grass, etc.) and use them to initialize the visual encoder, while the decoder is trained from random initialization.

The whole process is divided into three steps.

  • The first step is to pretrain the encoder with MAE on ImageNet;
  • the second step is to train the entire encoder-decoder offline on the ARC training set;
  • and the third step is to perform LoRA fine-tuning separately for each problem at test time.

During the fine-tuning stage, each problem has only 2 to 5 demonstration pairs, and the model uses these few demonstrations to try to "learn" the rules of that problem.

One detail worth noting: NAT-ARC directly used MAE's public checkpoint without spending a single penny on additional pretraining.

However, this checkpoint was originally designed for larger ImageNet images, while ARC grids are only 64×64 pixels—a large difference in size.

To adapt, NAT-ARC discarded the original patch embedding and positional encoding, keeping only the backbone weights, and replaced them with 2D RoPE as the positional encoding.

Simply put, NAT-ARC retained only the visual feature extraction ability that MAE learned on ImageNet, while discarding all of the spatial perception parts tied to ImageNet's image size, relearning them in a more flexible way.

But why would what was learned from looking at cat and dog photos help with abstract grid reasoning?

The paper offers a clue using attention visualization.

The researchers placed models with three initialization methods (no pretraining, ImageNet MAE pretraining, and ARC-style grid MAE pretraining) in front of the same ARC problem, observing their attention distributions before any ARC training had begun.

The result was clear: the randomly initialized model's attention was uniformly scattered with no focus, whereas the ImageNet-pretrained model could already distinguish foreground from background, with its attention automatically focusing on the meaningful pattern regions in the grid.

This model had never seen an ARC grid, yet the ability it learned on ImageNet to "recognize objects from their background" took effect naturally on ARC grids.

The researchers went further with a task-level breakdown.

They filtered out 15 ARC problems that showed significant improvement after pretraining, and after manual annotation found that these problems fell into two task categories.

Of these, 6 belonged to match-and-copy, where the model needs to identify matching objects, colors, or patterns, and then copy the corresponding structure into the output;

another 6 belonged to connected-component reasoning, where the model needs to identify, fill, or recolor spatially connected regions.

These two types of abilities happen to correspond to visual priors from natural images.

What was learned on ImageNet from looking at cats—"separating the cat from the grass background"—became in ARC "recognizing this connected colored region from the grid."

It turns out the underlying logic of the visual world and the abstract world shares the same set of representations in operations like object grouping, pattern matching, and region segmentation.

From cats, back to cats

NAT-ARC's final results on ARC-1 are as follows.

The best-performing single model uses the huge scale with 0.6B parameters, achieving a pass@2 score of 63.4±0.7%.

For ensembling, the researchers pooled together models trained separately under three different pretraining strategies (no pretraining, ImageNet MAE pretraining, and ARC-style grid MAE pretraining) and performed majority voting, achieving 70.2±0.6% pass@2 with about 2B total parameters.

As a reference, The ARChitects, who fine-tune specifically for ARC, achieved 71.6% using an 8B language model.

The visual route, with a quarter of the parameters, reached the gates of the specialized LLM systems.

Beyond the numbers, the most important finding of this paper is that pretraining broke through the scaling bottleneck of the visual route.

Without pretraining, larger models actually performed worse; with pretraining added, scaling began to deliver positive returns.

First, look at the training curves. In the offline training stage, models using ImageNet pretraining converge noticeably faster, and this holds across the base, large, and huge scales, with higher final accuracy as well.

But the most interesting finding hides in the scaling curves.

Without pretraining, performance still improves from base to large, but drops from large to huge.

The model gets bigger and results get worse — this is not a common phenomenon in deep learning, and it shows that the data for the ARC task is simply too scarce: what the increase in model capacity brings is not stronger generalization but overfitting.

With ImageNet pretraining added, the scaling curve becomes healthy: the base, large, and huge scales climb steadily upward, and the bigger the model, the better the performance.

One of the coolest experiments in the paper is a visual verification.

The researchers trained an autoencoder that maps ImageNet images onto a discrete grid representation of 60×60 cells with 10 colors — effectively 'translating' a cat photo into an ARC grid.

Then, they used NAT-ARC to apply an ARC transformation to this grid — a 180° rotation plus a blue border around it — and decoded it back into pixels.

The result: the cat was indeed rotated, and the border was indeed added. The model trained from cats returned, within the task, to the realm of cats.

This experiment turned 'natural images and ARC grids share the same representation space' from a statistical number into visual evidence.

The transformation rules learned on ARC grids can be applied directly to the latent representations of natural images, and vice versa.

The connection between abstract rules and visual representations is one these two worlds possessed all along.

Author Bios

The paper's first author, Xiaoman Delores Ding, a member of MIT CSAIL, comes from Kaiming He's group.

She is also the first author of VARC, and NAT-ARC effectively continues to push forward on the basis of her previous work.

The other authors include Keya Hu (胡珂雅), one of Kaiming He's first batch of female students, who completed her undergraduate studies atShanghai Jiao Tong University。

There are also Katelyn Gan and Victor Yin, also from MIT; both are undergraduates and serve as student researchers at CSAIL.

Corresponding author Kaiming He needs no introduction — he is the author of ResNet and MAE, and these two works alone can string together the main storyline of vision AI over the past decade.

Since joining MIT he has continuously produced foundational vision work, and this paper uses exactly his own MAE, proposed four years ago, as the pretraining tool — essentially using his own old weapon to open up a new path.

VARC has been accepted by CVPR 2026, and NAT-ARC is its direct follow-up.

Across two consecutive papers, this team has persisted in a pure-vision route to solving ARC; on a track dominated by LLMs, they are using experimental data to keep pushing the ceiling of the vision route.

Paper link:
https://eccv.ecva.net/virtual/2026/poster/5527

Original source

量子位

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original