arc agi 2 Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Claude Opus 5 Anthropic 90.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effortPublished Jul 24, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.4 100% confidence 100 percent, Full
2 GPT-5.5 OpenAI 85%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] VerifiedPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
85.0 100% confidence 100 percent, Full
3 GPT-5.4 Pro OpenAI 83.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] VerifiedPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
83.3 69% confidence 69 percent, Medium
4 Gemini 3.1 Pro Preview Google 77.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
77.1 100% confidence 100 percent, Full
5 GPT-5.4 OpenAI 73.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] VerifiedPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
73.3 100% confidence 100 percent, Full
6 Gemini 3.5 Flash Google 72.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
72.1 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.