gpqa diamond Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 0.5 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-6 Astra OpenAI 96%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Sep 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
96.0 100% confidence 100 percent, Full
2 GPT-6 Astra OpenAI 95.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.8 100% confidence 100 percent, Full
3 Claude Sonnet 5.5 Anthropic 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.6 93% confidence 93 percent, High
4 Fugu sakana 95.5%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
95.5 100% confidence 100 percent, Full
5 Fugu Ultra sakana 95.5%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
95.5 100% confidence 100 percent, Full
6 Gemini 3.8 Flash Google 95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.4 100% confidence 100 percent, Full
7 GPT-6.1 Sol OpenAI 95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.4 100% confidence 100 percent, Full
8 Gemini 3.7 Flash Google 94.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.8 100% confidence 100 percent, Full
9 GPT-5.4 Pro OpenAI 94.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Mar 20, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.6 69% confidence 69 percent, Medium
10 GPT-5.6 Sol OpenAI 94.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
94.6 100% confidence 100 percent, Full
11 Gemini 3.1 Pro Preview Google 94.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.4 100% confidence 100 percent, Full
12 GPT-5.4 Pro OpenAI 94.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
94.4 69% confidence 69 percent, Medium
13 GPT-6 Sol OpenAI 94.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.3 100% confidence 100 percent, Full
14 Gemini 3.1 Pro Preview Google 94.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
94.3 100% confidence 100 percent, Full
15 Claude Opus 4.7 Anthropic 94.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
94.2 100% confidence 100 percent, Full
16 Gemini 3.6 Flash Google 94.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.1 100% confidence 100 percent, Full
17 Claude Mythos 5 Anthropic 94.1%Anthropic Mythos 5 system-card claimsLab claim, not an independent evaluation. 198-question GPQA Diamond, average over 5 trials; June 9 release card section 8.8; immutable edition [variant] Mythos 5 system cardPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.1 5% confidence 5 percent, Low
18 Gemini 3.1 Pro Preview Google 94.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 20, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.1 100% confidence 100 percent, Full
19 Grok 4.6 xAI 94.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.0 100% confidence 100 percent, Full
20 Claude Opus 5 Anthropic 93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.9 100% confidence 100 percent, Full
21 GPT-5.5 OpenAI 93.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
93.6 100% confidence 100 percent, Full
22 GPT-5.6 Sol OpenAI 93.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.5 100% confidence 100 percent, Full
23 Grok 4.5 xAI 93.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 8, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.4 100% confidence 100 percent, Full
24 GPT-5.6 Terra OpenAI 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.3 100% confidence 100 percent, Full
25 GPT-5.4 OpenAI 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Mar 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.3 100% confidence 100 percent, Full
26 Grok 4.6 xAI 93.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.2 100% confidence 100 percent, Full
27 Kimi K3 Moonshot AI 93.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.1 100% confidence 100 percent, Full
28 Claude Opus 5 Anthropic 92.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.9 100% confidence 100 percent, Full
29 GPT-5.6 Terra OpenAI 92.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
92.9 100% confidence 100 percent, Full
30 Gemini 3.5 Flash Google 92.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.8 100% confidence 100 percent, Full
31 GPT-5.4 OpenAI 92.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
92.8 100% confidence 100 percent, Full
32 Grok 4.7 xAI 92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.7 100% confidence 100 percent, Full
33 Qwen3.8 Max Alibaba / Qwen 92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.7 85% confidence 85 percent, High
34 Gemini 3 Pro Preview Google 92.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 19, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.6 68% confidence 68 percent, Medium
35 Qwen3.8 Max Alibaba / Qwen 92.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
92.6 85% confidence 85 percent, High
36 Qwen3.8 Max Preview Alibaba / Qwen 92.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhighPublished Aug 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
92.6 6% confidence 6 percent, Low
37 Qwen3.7 Max Alibaba / Qwen 92.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
92.4 53% confidence 53 percent, Medium
38 GPT-5.6 Luna OpenAI 92.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
92.3 100% confidence 100 percent, Full
39 Qwen3.8 Max 0902 Alibaba / Qwen 92.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
92.3 40% confidence 40 percent, Low
40 Kimi K3 Moonshot AI 91.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.9 100% confidence 100 percent, Full
41 GLM-5.2 Z.ai 91.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 24, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.9 100% confidence 100 percent, Full
42 Qwen3.8 Flash Next Alibaba / Qwen 91.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
91.7 21% confidence 21 percent, Low
43 DeepSeek V4 Pro 0813 DeepSeek 91.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 18, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.7 69% confidence 69 percent, Medium
44 GPT-5.6 Luna OpenAI 91.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.6 100% confidence 100 percent, Full
45 GPT-5.2 OpenAI 91.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Dec 13, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.4 100% confidence 100 percent, Full
46 GLM-5.2 Z.ai 91.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jun 16, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
91.2 100% confidence 100 percent, Full
47 Claude Opus 4.8 Anthropic 91.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.0 100% confidence 100 percent, Full
48 DeepSeek V4 Flash 0731 DeepSeek 91.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.0 69% confidence 69 percent, Medium
49 Qwen3.7 Max Alibaba / Qwen 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.9 53% confidence 53 percent, Medium
50 DeepSeek V4 Pro DeepSeek 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.9 85% confidence 85 percent, High
51 MiniMax-M3 minimax 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.9 100% confidence 100 percent, Full
52 GLM-5.3 Z.ai 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 24, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.9 93% confidence 93 percent, High
53 DeepSeek V4.1 Flash DeepSeek 90.9%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.9 69% confidence 69 percent, Medium
54 Kimi K2.6 Moonshot AI 90.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 1, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.8 85% confidence 85 percent, High
55 GPT-5.5 OpenAI 90.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished May 5, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.7 100% confidence 100 percent, Full
56 Claude Opus 5.5 Anthropic 90.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.6 100% confidence 100 percent, Full
57 Claude Opus 4.6 Anthropic 90.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Feb 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.5 100% confidence 100 percent, Full
58 Claude Sonnet 5 Anthropic 90.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jul 1, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.5 93% confidence 93 percent, High
59 GPT-6 Luna OpenAI 90.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.5 100% confidence 100 percent, Full
60 Hy3 tencent 90.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] highest reasoning effort Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.4 88% confidence 88 percent, High
61 Claude Opus 4.7 Anthropic 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Apr 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.2 100% confidence 100 percent, Full
62 GLM-5.3-Flash Z.ai 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 26, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.2 100% confidence 100 percent, Full
63 DeepSeek V4 Pro DeepSeek 90.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.1 85% confidence 85 percent, High
64 GPT-5.4 OpenAI 89.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
89.9 100% confidence 100 percent, Full
65 GPT-5.6 Sol OpenAI 89.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
89.9 100% confidence 100 percent, Full
66 GLM-5.1 Z.ai 89.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
89.9 64% confidence 64 percent, Medium
67 DeepSeek V4 Pro DeepSeek 89.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 16, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
89.6 85% confidence 85 percent, High
68 Gemini 3 Flash Preview Google 89.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
89.4 93% confidence 93 percent, High
69 Grok 4.20 (Reasoning) xAI 89.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
89.3 80% confidence 80 percent, High
70 Qwen3.8 27B Alibaba / Qwen 89.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
89.2 69% confidence 69 percent, Medium
71 LongCat-2.0 meituan 88.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jun 30, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.9 100% confidence 100 percent, Full
72 Gemini 3.5 Flash Google 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.9 100% confidence 100 percent, Full
73 GPT-5.4 OpenAI 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.9 100% confidence 100 percent, Full
74 Grok 4.3 xAI 88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.8 80% confidence 80 percent, High
75 Claude Opus 4.6 Anthropic 88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Feb 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.8 100% confidence 100 percent, Full
76 Inkling Small thinkingmachines 88.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.5 100% confidence 100 percent, Full
77 Qwen3.6 Plus Alibaba / Qwen 88.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.4 80% confidence 80 percent, High
78 Claude Opus 4.6 Anthropic 88.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.4 100% confidence 100 percent, Full
79 Claude Opus 4.8 Anthropic 88.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.4 100% confidence 100 percent, Full
80 Inkling thinkingmachines 88.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 5, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.3 100% confidence 100 percent, Full
81 GPT-5.2 OpenAI 88.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Dec 11, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
88.2 100% confidence 100 percent, Full
82 DeepSeek V4 Flash DeepSeek 88.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.1 53% confidence 53 percent, Medium
83 GPT-5.4 mini OpenAI 88%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhighPublished Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.0 100% confidence 100 percent, Full
84 Qwen3.7 Plus Alibaba / Qwen 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.9 64% confidence 64 percent, Medium
85 Claude Opus 5 Anthropic 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.9 100% confidence 100 percent, Full
86 Kimi K2.7 Code Moonshot AI 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.9 48% confidence 48 percent, Low
87 GPT-5.2 OpenAI 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Dec 11, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.9 100% confidence 100 percent, Full
88 GLM-5.2 Z.ai 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.9 100% confidence 100 percent, Full
89 GLM-5 Z.ai 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.8 90% confidence 90 percent, High
90 GPT-5.1 OpenAI 87.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Nov 13, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.6 85% confidence 85 percent, High
91 Qwen3.6 Max Preview Alibaba / Qwen 87.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.4 64% confidence 64 percent, Medium
92 Claude Sonnet 4.6 Anthropic 87.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Feb 20, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.4 100% confidence 100 percent, Full
93 GPT-5.6 Terra OpenAI 87.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.4 100% confidence 100 percent, Full
94 Nemotron 3 Ultra 550B A55B nvidia 87%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 4, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
87.0 100% confidence 100 percent, Full
95 GPT-5.4 mini OpenAI 86.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.9 100% confidence 100 percent, Full
96 MiMo-V2.5-Pro xiaomi 86.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 22, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
86.6 88% confidence 88 percent, High
97 Qwen3.5 397B-A17B Alibaba / Qwen 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.4 69% confidence 69 percent, Medium
98 Claude Opus 4.7 Anthropic 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.4 100% confidence 100 percent, Full
99 Gemini 3.5 Flash Google 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.4 100% confidence 100 percent, Full
100 Gemini 3.6 Flash Google 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.4 100% confidence 100 percent, Full
101 GPT-5 OpenAI 86.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.2 100% confidence 100 percent, Full
102 Claude Opus 4.5 Anthropic 86.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Nov 24, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.0 100% confidence 100 percent, Full
103 Qwen3.5 397B-A17B Alibaba / Qwen 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.9 69% confidence 69 percent, Medium
104 Qwen3.6 27B Alibaba / Qwen 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.9 53% confidence 53 percent, Medium
105 Claude Fable 5 Anthropic 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.9 100% confidence 100 percent, Full
106 Gemini 3.6 Flash Google 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.9 100% confidence 100 percent, Full
107 Claude Opus 4.5 Anthropic 85.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Nov 25, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.5 100% confidence 100 percent, Full
108 Claude Opus 4.8 Anthropic 85.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.4 100% confidence 100 percent, Full
109 Nemotron 3 Ultra 550B A55B nvidia 85.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.4 100% confidence 100 percent, Full
110 GPT-5 OpenAI 85.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.4 100% confidence 100 percent, Full
111 Gemini 2.5 Pro Google 85.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.3 90% confidence 90 percent, High
112 GPT-5.1 OpenAI 85.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Nov 17, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.0 85% confidence 85 percent, High
113 Qwen3.5 Plus Alibaba / Qwen 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.8 32% confidence 32 percent, Low
114 Qwen3.6 27B Alibaba / Qwen 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.8 53% confidence 53 percent, Medium
115 Qwen3.6 35B-A3B Alibaba / Qwen 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.8 37% confidence 37 percent, Low
116 Kimi K3 Moonshot AI 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.8 100% confidence 100 percent, Full
117 GPT-5.4 OpenAI 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.8 100% confidence 100 percent, Full
118 Kimi K2 Thinking Turbo Moonshot AI 84.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 11, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.2 64% confidence 64 percent, Medium
119 Qwen3.6 35B-A3B Alibaba / Qwen 83.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.8 37% confidence 37 percent, Low
120 GPT-5.4 mini OpenAI 83.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.6 100% confidence 100 percent, Full
121 Muse Glimmer 30B Meta 83.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] AAPublished Aug 10, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
83.5 6% confidence 6 percent, Low
122 Qwen3.5 35B-A3B Alibaba / Qwen 83.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.5 64% confidence 64 percent, Medium
123 DeepSeek Reasoner DeepSeek 83.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Dec 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.4 32% confidence 32 percent, Low
124 Qwen3.6 Flash Alibaba / Qwen 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.3 32% confidence 32 percent, Low
125 Claude Fable 5 Anthropic 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.3 100% confidence 100 percent, Full
126 Claude Sonnet 4.6 Anthropic 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.3 100% confidence 100 percent, Full
127 Claude Sonnet 4.6 Anthropic 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.3 100% confidence 100 percent, Full
128 Gemini 3.5 Flash Lite Google 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.3 100% confidence 100 percent, Full
129 GLM-4.7 Z.ai 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 29, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.3 69% confidence 69 percent, Medium
130 Gemini 3 Flash Preview Google 83.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Dec 17, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.2 93% confidence 93 percent, High
131 GPT-5.6 Sol OpenAI 82.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.8 100% confidence 100 percent, Full
132 GPT-5.4 nano OpenAI 82.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhighPublished Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
82.8 100% confidence 100 percent, Full
133 GPT-5.2 OpenAI 82.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Dec 11, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.7 100% confidence 100 percent, Full
134 GPT-5.5 Instant OpenAI 82.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.5 64% confidence 64 percent, Medium
135 Qwen3.5 Flash Alibaba / Qwen 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.3 64% confidence 64 percent, Medium
136 Qwen3.7 Flash Alibaba / Qwen 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.3 32% confidence 32 percent, Low
137 Claude Sonnet 4.5 Anthropic 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished Oct 28, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.3 80% confidence 80 percent, High
138 GPT-5.6 Luna OpenAI 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.3 100% confidence 100 percent, Full
139 Qwen3.7 Plus Alibaba / Qwen 81.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.8 64% confidence 64 percent, Medium
140 Gemini 3.1 Flash Lite Google 81.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.8 32% confidence 32 percent, Low
141 o3 OpenAI 81.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.8 87% confidence 87 percent, High
142 Claude Sonnet 4.5 Anthropic 81.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 21, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.7 80% confidence 80 percent, High
143 MiniMax-M3 minimax 81.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.3 100% confidence 100 percent, Full
144 Qwen3.5 35B-A3B Alibaba / Qwen 81.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.2 64% confidence 64 percent, Medium
145 Qwen3.7 Flash Alibaba / Qwen 80.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
80.8 32% confidence 32 percent, Low
146 o3 OpenAI 80.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
80.8 87% confidence 87 percent, High
147 Claude Opus 4.5 Anthropic 80.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 24, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
80.7 100% confidence 100 percent, Full
148 Claude Sonnet 5 Anthropic 80.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
80.3 93% confidence 93 percent, High
149 o3 OpenAI 79.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
79.8 87% confidence 87 percent, High
150 o4-mini OpenAI 79.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
79.6 87% confidence 87 percent, High
151 Qwen3.5 9B Alibaba / Qwen 79.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
79.0 32% confidence 32 percent, Low
152 Qwen3.5 9B Alibaba / Qwen 78.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.9 32% confidence 32 percent, Low
153 Claude Fable 5 Anthropic 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.8 100% confidence 100 percent, Full
154 Claude Sonnet 4.5 Anthropic 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Oct 28, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.8 80% confidence 80 percent, High
155 Claude Sonnet 4.6 Anthropic 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.8 100% confidence 100 percent, Full
156 Claude Sonnet 3.7 Anthropic 78.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished May 26, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.5 100% confidence 100 percent, Full
157 GPT-5.4 nano OpenAI 78.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.5 100% confidence 100 percent, Full
158 Claude Sonnet 4 Anthropic 78.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.3 100% confidence 100 percent, Full
159 Claude Sonnet 4 Anthropic 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished May 26, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
77.8 100% confidence 100 percent, Full
160 o4-mini OpenAI 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
77.8 87% confidence 87 percent, High
161 Claude Opus 4.1 Anthropic 77.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Aug 5, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
77.3 80% confidence 80 percent, High
162 GPT-5.5 OpenAI 77.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
77.3 100% confidence 100 percent, Full
163 GPT-5.6 Terra OpenAI 77.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
77.3 100% confidence 100 percent, Full
164 Claude Sonnet 3.7 Anthropic 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
76.8 100% confidence 100 percent, Full
165 Claude Sonnet 3.7 Anthropic 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 10, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
76.8 100% confidence 100 percent, Full
166 Claude Opus 4.1 Anthropic 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 27KPublished Aug 5, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
76.8 80% confidence 80 percent, High
167 DeepSeek-R1 DeepSeek 76.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
76.3 88% confidence 88 percent, High
168 Claude Opus 4 Anthropic 76.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
76.3 100% confidence 100 percent, Full
169 Claude Sonnet 4 Anthropic 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.8 100% confidence 100 percent, Full
170 Gemini 3.5 Flash Lite Google 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.8 100% confidence 100 percent, Full
171 Gemma 4 31B IT Google 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.8 64% confidence 64 percent, Medium
172 GPT OSS 120B OpenAI 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Dec 11, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.8 79% confidence 79 percent, Medium
173 Nemotron 3.5 Lightning 30B A3B nvidia 75.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; no tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
75.4 100% confidence 100 percent, Full
174 o4-mini OpenAI 75.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.3 87% confidence 87 percent, High
175 GPT-5 Mini OpenAI 75%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.0 99% confidence 99 percent, High
176 GPT-5.4 OpenAI 74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Jul 15, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
74.7 100% confidence 100 percent, Full
177 Gemini 3.1 Flash Lite Google 74.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
74.2 32% confidence 32 percent, Low
178 Gemini 3.5 Flash Lite Google 74.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
74.2 100% confidence 100 percent, Full
179 Claude Sonnet 4.5 Anthropic 73.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Sep 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.7 80% confidence 80 percent, High
180 Gemini 3.1 Flash Lite Google 73.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.7 32% confidence 32 percent, Low
181 Claude Opus 4.1 Anthropic 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 5, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.2 80% confidence 80 percent, High
182 DeepSeek V4 Pro DeepSeek 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.2 85% confidence 85 percent, High
183 Gemma 4 26B A4B IT Google 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.2 64% confidence 64 percent, Medium
184 GPT-5.2 OpenAI 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.2 100% confidence 100 percent, Full
185 Qwen3 Max Alibaba / Qwen 72.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 6, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
72.6 69% confidence 69 percent, Medium
186 GPT-5.4 nano OpenAI 72.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
72.2 100% confidence 100 percent, Full
187 GPT-5 OpenAI 71.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 20, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.7 100% confidence 100 percent, Full
188 GPT-5 Mini OpenAI 71.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.7 99% confidence 99 percent, High
189 GPT-5 Mini OpenAI 71.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.7 99% confidence 99 percent, High
190 Claude Haiku 4.5 Anthropic 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.2 64% confidence 64 percent, Medium
191 DeepSeek Chat DeepSeek 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.2 32% confidence 32 percent, Low
192 GLM-5.2 Z.ai 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.2 100% confidence 100 percent, Full
193 Qwen3 235B-A22B Alibaba / Qwen 70.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 3, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
70.7 76% confidence 76 percent, Medium
194 GPT-5 Nano OpenAI 69.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
69.4 85% confidence 85 percent, High
195 Claude Opus 4 Anthropic 69.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
69.2 100% confidence 100 percent, Full
196 DeepSeek V3 0324 DeepSeek 67.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 1, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
67.6 74% confidence 74 percent, Medium
197 GPT-5 Nano OpenAI 67.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
67.4 85% confidence 85 percent, High
198 Llama 4 Maverick 17B Instruct Meta 67.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
67.0 78% confidence 78 percent, Medium
199 GPT-4.1 OpenAI 66.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
66.9 100% confidence 100 percent, Full
200 Claude Sonnet 4 Anthropic 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
66.7 100% confidence 100 percent, Full
201 GPT-5.1 OpenAI 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
66.7 85% confidence 85 percent, High
202 Claude Sonnet 3.7 Anthropic 66.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 24, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
66.0 100% confidence 100 percent, Full
203 GPT-4.1 mini OpenAI 65.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
65.8 87% confidence 87 percent, High
204 Qwen3 32B Alibaba / Qwen 65.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
65.7 67% confidence 67 percent, Medium
205 QwQ Plus Alibaba / Qwen 65.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 11, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
65.4 32% confidence 32 percent, Low
206 QwQ 32B Alibaba / Qwen 65.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
65.3 67% confidence 67 percent, Medium
207 DeepSeek-R1-Distill-Qwen-32B DeepSeek 64.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
64.1 32% confidence 32 percent, Low
208 GPT-5.4 mini OpenAI 64.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
64.1 100% confidence 100 percent, Full
209 GPT-5.6 Luna OpenAI 63.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
63.6 100% confidence 100 percent, Full
210 Qwen3 30B A3B Alibaba / Qwen 61.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
61.7 64% confidence 64 percent, Medium
211 GPT OSS 20B OpenAI 60.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
60.8 64% confidence 64 percent, Medium
212 GLM-4.7-Flash Z.ai 60.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
60.5 69% confidence 69 percent, Medium
213 Claude Haiku 4.5 Anthropic 60.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
60.5 64% confidence 64 percent, Medium
214 Mistral Medium 3 mistral 59.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
59.5 85% confidence 85 percent, High
215 GPT-5 Nano OpenAI 57.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
57.6 85% confidence 85 percent, High
216 DeepSeek-V3 DeepSeek 56.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.5 93% confidence 93 percent, High
217 Magistral Small mistral 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.1 48% confidence 48 percent, Low
218 GPT-5.4 nano OpenAI 55.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
55.6 100% confidence 100 percent, Full
219 Claude Sonnet 3.5 v2 Anthropic 55.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
55.3 100% confidence 100 percent, Full
220 Qwen3 32B Alibaba / Qwen 54.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
54.1 67% confidence 67 percent, Medium
221 GPT OSS 20B OpenAI 53.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
53.2 64% confidence 64 percent, Medium
222 Llama 4 Scout 17B Instruct Meta 51.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
51.8 64% confidence 64 percent, Medium
223 Mistral Large 2.1 mistral 51.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
51.3 100% confidence 100 percent, Full
224 Qwen3 30B A3B Alibaba / Qwen 50.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
50.4 64% confidence 64 percent, Medium
225 GPT-4o (2024-08-06) OpenAI 49.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
49.2 100% confidence 100 percent, Full
226 GPT-4.1 nano OpenAI 48.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
48.9 82% confidence 82 percent, High
227 GPT-4o (2024-05-13) OpenAI 48.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
48.9 93% confidence 93 percent, High
228 GPT-5 Nano OpenAI 48.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
48.5 85% confidence 85 percent, High
229 GPT-4o (2024-11-20) OpenAI 47.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 5, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
47.9 61% confidence 61 percent, Medium
230 Gemma 3 27B IT Google 47.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
47.7 74% confidence 74 percent, Medium
231 Magistral Small 1.2 mistral 47.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
47.6 32% confidence 32 percent, Low
232 Llama-3.3-70B-Instruct Meta 47.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
47.4 100% confidence 100 percent, Full
233 GPT OSS 20B OpenAI 46.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
46.0 64% confidence 64 percent, Medium
234 GLM-4.7-Flash Z.ai 45.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
45.1 69% confidence 69 percent, Medium
235 Llama-3.1-70B-Instruct Meta 44.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
44.2 93% confidence 93 percent, High
236 Qwen Turbo Alibaba / Qwen 41.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
41.8 47% confidence 47 percent, Low
237 Gemma 3 12B IT Google 39.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
39.5 64% confidence 64 percent, Medium
238 Claude Haiku 3.5 Anthropic 38.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 12, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
38.1 100% confidence 100 percent, Full
239 GPT-4o mini OpenAI 37.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
37.7 100% confidence 100 percent, Full
240 Claude Haiku 3 Anthropic 36.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
36.3 93% confidence 93 percent, High
241 Llama-3.1-8B-Instruct Meta 27.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
27.0 93% confidence 93 percent, High
242 Llama-3.2-1B Meta 23.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
23.9 93% confidence 93 percent, High
243 Gemma 3 4B IT Google 23.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
23.2 64% confidence 64 percent, Medium

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.