Kaiming He's MIT group: ImageNet pretraining lifts a pure-vision ARC solver (Nat-ARC) to 63.4% on ARC-AGI-1, 70.2% ensembled
"Natural Image Pretraining Improves Abstract Reasoning" (Ding, Hu, Gan, Yin, Kaiming He; MIT; ECCV 2026) introduces Nat-ARC. It extends the vision-only VARC pipeline by initializing its ViT encoder from Masked Autoencoders pretrained on ImageNet. The best single model reaches 63.4 ± 0.7% pass@2 on ARC-1, and an ensemble reaches 70.2 ± 0.6%, with no language model. Chinese media covered it on Oct 1, 2026.
Key facts
- Best ImageNet-MAE-pretrained single model (huge, ~0.6B parameters per QbitAI): 63.4 ± 0.7% pass@2 on ARC-1
- Prediction-level ensemble of no pretraining, ImageNet MAE pretraining and ARC-style grid pretraining: 70.2 ± 0.6% pass@2 on ARC-1, 'using a purely vision-based method'
- ImageNet MAE pretraining improves over matched randomly initialized baselines at base, large and huge scales
- Predecessor: VARC ('ARC Is a Vision Problem!', arXiv 2511.14761, Nov 2025), 60.4% on ARC-1 with an 18M-parameter ViT trained from scratch
- Tasks helped most involve visual matching, copying and connected-component reasoning (authors' qualitative analysis)
- QbitAI's comparison: LoopViT 65.8%, Loop-OWM 68.5%, and the LLM-based ARChitects (8B) 71.6% (as reported by QbitAI, not checked against the paper's tables)
What happened
Kaiming He's group at MIT asked whether visual knowledge from natural photos transfers to ARC's abstract colored-grid puzzles. They took the VARC pipeline, which treats each ARC task as image-to-image translation with offline training plus per-task test-time training. They replaced its random initialization with an encoder pretrained as a Masked Autoencoder on ImageNet. Pretraining helped at every model scale, and ensembling differently pretrained models reached 70.2% pass@2 on ARC-1 without any language model.
Why it matters
Almost all high ARC-AGI-1 scores come from LLM pipelines. Nat-ARC shows that a small vision-only model with generic visual pretraining gets close, which supports the view that much of ARC is perception plus few-shot adaptation. ARC-1 is now largely saturated by frontier LLMs, and the paper reports no ARC-AGI-2 or ARC-AGI-3 results.
Changelog
- 2026-10-02: created (leads: Sina AI hourly report Oct 2, QbitAI)
Related events
Sources (4)
- paperECCV 2026 poster: Natural Image Pretraining Improves Abstract Reasoning
- paperPaper PDF (ECCV 2026)
- pressQbitAI: Kaiming He's team learns ARC from cat pictures (Chinese)
- paperarXiv 2511.14761: ARC Is a Vision Problem! (VARC)
id: 2026-10-01-nat-arc-natural-image-pretraining-arc · updated 2026-10-02 · open in the interactive timeline