gpqa diamond Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-6 Astra OpenAI | 96%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 96.0 | 100% confidence 100 percent, Full |
| 2 | GPT-6 Astra OpenAI | 95.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.8 | 100% confidence 100 percent, Full |
| 3 | Claude Sonnet 5.5 Anthropic | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 93% confidence 93 percent, High |
| 4 | Fugu sakana | 95.5%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 5 | Fugu Ultra sakana | 95.5%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 6 | Gemini 3.8 Flash Google | 95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.4 | 100% confidence 100 percent, Full |
| 7 | GPT-6.1 Sol OpenAI | 95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.4 | 100% confidence 100 percent, Full |
| 8 | Gemini 3.7 Flash Google | 94.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.8 | 100% confidence 100 percent, Full |
| 9 | GPT-5.4 Pro OpenAI | 94.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Mar 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.6 | 69% confidence 69 percent, Medium |
| 10 | GPT-5.6 Sol OpenAI | 94.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 94.6 | 100% confidence 100 percent, Full |
| 11 | Gemini 3.1 Pro Preview Google | 94.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.4 | 100% confidence 100 percent, Full |
| 12 | GPT-5.4 Pro OpenAI | 94.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 94.4 | 69% confidence 69 percent, Medium |
| 13 | GPT-6 Sol OpenAI | 94.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.3 | 100% confidence 100 percent, Full |
| 14 | Gemini 3.1 Pro Preview Google | 94.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 94.3 | 100% confidence 100 percent, Full |
| 15 | Claude Opus 4.7 Anthropic | 94.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 94.2 | 100% confidence 100 percent, Full |
| 16 | Gemini 3.6 Flash Google | 94.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.1 | 100% confidence 100 percent, Full |
| 17 | Claude Mythos 5 Anthropic | 94.1%Anthropic Mythos 5 system-card claimsLab claim, not an independent evaluation. 198-question GPQA Diamond, average over 5 trials; June 9 release card section 8.8; immutable edition [variant] Mythos 5 system cardPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.1 | 5% confidence 5 percent, Low |
| 18 | Gemini 3.1 Pro Preview Google | 94.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.1 | 100% confidence 100 percent, Full |
| 19 | Grok 4.6 xAI | 94.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.0 | 100% confidence 100 percent, Full |
| 20 | Claude Opus 5 Anthropic | 93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.9 | 100% confidence 100 percent, Full |
| 21 | GPT-5.5 OpenAI | 93.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 93.6 | 100% confidence 100 percent, Full |
| 22 | GPT-5.6 Sol OpenAI | 93.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.5 | 100% confidence 100 percent, Full |
| 23 | Grok 4.5 xAI | 93.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 8, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.4 | 100% confidence 100 percent, Full |
| 24 | GPT-5.6 Terra OpenAI | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 25 | GPT-5.4 OpenAI | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Mar 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 26 | Grok 4.6 xAI | 93.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.2 | 100% confidence 100 percent, Full |
| 27 | Kimi K3 Moonshot AI | 93.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.1 | 100% confidence 100 percent, Full |
| 28 | Claude Opus 5 Anthropic | 92.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.9 | 100% confidence 100 percent, Full |
| 29 | GPT-5.6 Terra OpenAI | 92.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 92.9 | 100% confidence 100 percent, Full |
| 30 | Gemini 3.5 Flash Google | 92.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.8 | 100% confidence 100 percent, Full |
| 31 | GPT-5.4 OpenAI | 92.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 92.8 | 100% confidence 100 percent, Full |
| 32 | Grok 4.7 xAI | 92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.7 | 100% confidence 100 percent, Full |
| 33 | Qwen3.8 Max Alibaba / Qwen | 92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.7 | 85% confidence 85 percent, High |
| 34 | Gemini 3 Pro Preview Google | 92.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 19, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.6 | 68% confidence 68 percent, Medium |
| 35 | Qwen3.8 Max Alibaba / Qwen | 92.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 92.6 | 85% confidence 85 percent, High |
| 36 | Qwen3.8 Max Preview Alibaba / Qwen | 92.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhighPublished Aug 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 92.6 | 6% confidence 6 percent, Low |
| 37 | Qwen3.7 Max Alibaba / Qwen | 92.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 92.4 | 53% confidence 53 percent, Medium |
| 38 | GPT-5.6 Luna OpenAI | 92.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 92.3 | 100% confidence 100 percent, Full |
| 39 | Qwen3.8 Max 0902 Alibaba / Qwen | 92.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.3 | 40% confidence 40 percent, Low |
| 40 | Kimi K3 Moonshot AI | 91.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.9 | 100% confidence 100 percent, Full |
| 41 | GLM-5.2 Z.ai | 91.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.9 | 100% confidence 100 percent, Full |
| 42 | Qwen3.8 Flash Next Alibaba / Qwen | 91.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 91.7 | 21% confidence 21 percent, Low |
| 43 | DeepSeek V4 Pro 0813 DeepSeek | 91.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 18, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.7 | 69% confidence 69 percent, Medium |
| 44 | GPT-5.6 Luna OpenAI | 91.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.6 | 100% confidence 100 percent, Full |
| 45 | GPT-5.2 OpenAI | 91.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Dec 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.4 | 100% confidence 100 percent, Full |
| 46 | GLM-5.2 Z.ai | 91.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jun 16, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 91.2 | 100% confidence 100 percent, Full |
| 47 | Claude Opus 4.8 Anthropic | 91.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.0 | 100% confidence 100 percent, Full |
| 48 | DeepSeek V4 Flash 0731 DeepSeek | 91.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.0 | 69% confidence 69 percent, Medium |
| 49 | Qwen3.7 Max Alibaba / Qwen | 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.9 | 53% confidence 53 percent, Medium |
| 50 | DeepSeek V4 Pro DeepSeek | 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.9 | 85% confidence 85 percent, High |
| 51 | MiniMax-M3 minimax | 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.9 | 100% confidence 100 percent, Full |
| 52 | GLM-5.3 Z.ai | 90.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.9 | 93% confidence 93 percent, High |
| 53 | DeepSeek V4.1 Flash DeepSeek | 90.9%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.9 | 69% confidence 69 percent, Medium |
| 54 | Kimi K2.6 Moonshot AI | 90.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 1, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.8 | 85% confidence 85 percent, High |
| 55 | GPT-5.5 OpenAI | 90.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished May 5, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.7 | 100% confidence 100 percent, Full |
| 56 | Claude Opus 5.5 Anthropic | 90.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.6 | 100% confidence 100 percent, Full |
| 57 | Claude Opus 4.6 Anthropic | 90.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Feb 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.5 | 100% confidence 100 percent, Full |
| 58 | Claude Sonnet 5 Anthropic | 90.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jul 1, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.5 | 93% confidence 93 percent, High |
| 59 | GPT-6 Luna OpenAI | 90.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.5 | 100% confidence 100 percent, Full |
| 60 | Hy3 tencent | 90.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] highest reasoning effort
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.4 | 88% confidence 88 percent, High |
| 61 | Claude Opus 4.7 Anthropic | 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Apr 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.2 | 100% confidence 100 percent, Full |
| 62 | GLM-5.3-Flash Z.ai | 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 26, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.2 | 100% confidence 100 percent, Full |
| 63 | DeepSeek V4 Pro DeepSeek | 90.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.1 | 85% confidence 85 percent, High |
| 64 | GPT-5.4 OpenAI | 89.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.9 | 100% confidence 100 percent, Full |
| 65 | GPT-5.6 Sol OpenAI | 89.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.9 | 100% confidence 100 percent, Full |
| 66 | GLM-5.1 Z.ai | 89.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.9 | 64% confidence 64 percent, Medium |
| 67 | DeepSeek V4 Pro DeepSeek | 89.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.6 | 85% confidence 85 percent, High |
| 68 | Gemini 3 Flash Preview Google | 89.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.4 | 93% confidence 93 percent, High |
| 69 | Grok 4.20 (Reasoning) xAI | 89.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.3 | 80% confidence 80 percent, High |
| 70 | Qwen3.8 27B Alibaba / Qwen | 89.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 89.2 | 69% confidence 69 percent, Medium |
| 71 | LongCat-2.0 meituan | 88.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jun 30, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 72 | Gemini 3.5 Flash Google | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 73 | GPT-5.4 OpenAI | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 74 | Grok 4.3 xAI | 88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.8 | 80% confidence 80 percent, High |
| 75 | Claude Opus 4.6 Anthropic | 88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Feb 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.8 | 100% confidence 100 percent, Full |
| 76 | Inkling Small thinkingmachines | 88.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.5 | 100% confidence 100 percent, Full |
| 77 | Qwen3.6 Plus Alibaba / Qwen | 88.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.4 | 80% confidence 80 percent, High |
| 78 | Claude Opus 4.6 Anthropic | 88.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.4 | 100% confidence 100 percent, Full |
| 79 | Claude Opus 4.8 Anthropic | 88.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.4 | 100% confidence 100 percent, Full |
| 80 | Inkling thinkingmachines | 88.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 5, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.3 | 100% confidence 100 percent, Full |
| 81 | GPT-5.2 OpenAI | 88.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.2 | 100% confidence 100 percent, Full |
| 82 | DeepSeek V4 Flash DeepSeek | 88.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.1 | 53% confidence 53 percent, Medium |
| 83 | GPT-5.4 mini OpenAI | 88%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhighPublished Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.0 | 100% confidence 100 percent, Full |
| 84 | Qwen3.7 Plus Alibaba / Qwen | 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.9 | 64% confidence 64 percent, Medium |
| 85 | Claude Opus 5 Anthropic | 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.9 | 100% confidence 100 percent, Full |
| 86 | Kimi K2.7 Code Moonshot AI | 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.9 | 48% confidence 48 percent, Low |
| 87 | GPT-5.2 OpenAI | 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.9 | 100% confidence 100 percent, Full |
| 88 | GLM-5.2 Z.ai | 87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.9 | 100% confidence 100 percent, Full |
| 89 | GLM-5 Z.ai | 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.8 | 90% confidence 90 percent, High |
| 90 | GPT-5.1 OpenAI | 87.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Nov 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.6 | 85% confidence 85 percent, High |
| 91 | Qwen3.6 Max Preview Alibaba / Qwen | 87.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.4 | 64% confidence 64 percent, Medium |
| 92 | Claude Sonnet 4.6 Anthropic | 87.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Feb 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.4 | 100% confidence 100 percent, Full |
| 93 | GPT-5.6 Terra OpenAI | 87.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.4 | 100% confidence 100 percent, Full |
| 94 | Nemotron 3 Ultra 550B A55B nvidia | 87%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 4, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 87.0 | 100% confidence 100 percent, Full |
| 95 | GPT-5.4 mini OpenAI | 86.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.9 | 100% confidence 100 percent, Full |
| 96 | MiMo-V2.5-Pro xiaomi | 86.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Apr 22, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 86.6 | 88% confidence 88 percent, High |
| 97 | Qwen3.5 397B-A17B Alibaba / Qwen | 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.4 | 69% confidence 69 percent, Medium |
| 98 | Claude Opus 4.7 Anthropic | 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.4 | 100% confidence 100 percent, Full |
| 99 | Gemini 3.5 Flash Google | 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.4 | 100% confidence 100 percent, Full |
| 100 | Gemini 3.6 Flash Google | 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.4 | 100% confidence 100 percent, Full |
| 101 | GPT-5 OpenAI | 86.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.2 | 100% confidence 100 percent, Full |
| 102 | Claude Opus 4.5 Anthropic | 86.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Nov 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.0 | 100% confidence 100 percent, Full |
| 103 | Qwen3.5 397B-A17B Alibaba / Qwen | 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.9 | 69% confidence 69 percent, Medium |
| 104 | Qwen3.6 27B Alibaba / Qwen | 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.9 | 53% confidence 53 percent, Medium |
| 105 | Claude Fable 5 Anthropic | 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.9 | 100% confidence 100 percent, Full |
| 106 | Gemini 3.6 Flash Google | 85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.9 | 100% confidence 100 percent, Full |
| 107 | Claude Opus 4.5 Anthropic | 85.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Nov 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.5 | 100% confidence 100 percent, Full |
| 108 | Claude Opus 4.8 Anthropic | 85.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.4 | 100% confidence 100 percent, Full |
| 109 | Nemotron 3 Ultra 550B A55B nvidia | 85.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.4 | 100% confidence 100 percent, Full |
| 110 | GPT-5 OpenAI | 85.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.4 | 100% confidence 100 percent, Full |
| 111 | Gemini 2.5 Pro Google | 85.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.3 | 90% confidence 90 percent, High |
| 112 | GPT-5.1 OpenAI | 85.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Nov 17, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.0 | 85% confidence 85 percent, High |
| 113 | Qwen3.5 Plus Alibaba / Qwen | 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.8 | 32% confidence 32 percent, Low |
| 114 | Qwen3.6 27B Alibaba / Qwen | 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.8 | 53% confidence 53 percent, Medium |
| 115 | Qwen3.6 35B-A3B Alibaba / Qwen | 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.8 | 37% confidence 37 percent, Low |
| 116 | Kimi K3 Moonshot AI | 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.8 | 100% confidence 100 percent, Full |
| 117 | GPT-5.4 OpenAI | 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.8 | 100% confidence 100 percent, Full |
| 118 | Kimi K2 Thinking Turbo Moonshot AI | 84.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.2 | 64% confidence 64 percent, Medium |
| 119 | Qwen3.6 35B-A3B Alibaba / Qwen | 83.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.8 | 37% confidence 37 percent, Low |
| 120 | GPT-5.4 mini OpenAI | 83.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.6 | 100% confidence 100 percent, Full |
| 121 | Muse Glimmer 30B Meta | 83.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] AAPublished Aug 10, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 83.5 | 6% confidence 6 percent, Low |
| 122 | Qwen3.5 35B-A3B Alibaba / Qwen | 83.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.5 | 64% confidence 64 percent, Medium |
| 123 | DeepSeek Reasoner DeepSeek | 83.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Dec 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.4 | 32% confidence 32 percent, Low |
| 124 | Qwen3.6 Flash Alibaba / Qwen | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 32% confidence 32 percent, Low |
| 125 | Claude Fable 5 Anthropic | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 100% confidence 100 percent, Full |
| 126 | Claude Sonnet 4.6 Anthropic | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 100% confidence 100 percent, Full |
| 127 | Claude Sonnet 4.6 Anthropic | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 100% confidence 100 percent, Full |
| 128 | Gemini 3.5 Flash Lite Google | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 100% confidence 100 percent, Full |
| 129 | GLM-4.7 Z.ai | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 69% confidence 69 percent, Medium |
| 130 | Gemini 3 Flash Preview Google | 83.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Dec 17, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.2 | 93% confidence 93 percent, High |
| 131 | GPT-5.6 Sol OpenAI | 82.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.8 | 100% confidence 100 percent, Full |
| 132 | GPT-5.4 nano OpenAI | 82.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhighPublished Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 82.8 | 100% confidence 100 percent, Full |
| 133 | GPT-5.2 OpenAI | 82.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.7 | 100% confidence 100 percent, Full |
| 134 | GPT-5.5 Instant OpenAI | 82.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.5 | 64% confidence 64 percent, Medium |
| 135 | Qwen3.5 Flash Alibaba / Qwen | 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.3 | 64% confidence 64 percent, Medium |
| 136 | Qwen3.7 Flash Alibaba / Qwen | 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.3 | 32% confidence 32 percent, Low |
| 137 | Claude Sonnet 4.5 Anthropic | 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished Oct 28, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.3 | 80% confidence 80 percent, High |
| 138 | GPT-5.6 Luna OpenAI | 82.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.3 | 100% confidence 100 percent, Full |
| 139 | Qwen3.7 Plus Alibaba / Qwen | 81.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.8 | 64% confidence 64 percent, Medium |
| 140 | Gemini 3.1 Flash Lite Google | 81.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.8 | 32% confidence 32 percent, Low |
| 141 | o3 OpenAI | 81.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.8 | 87% confidence 87 percent, High |
| 142 | Claude Sonnet 4.5 Anthropic | 81.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 21, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.7 | 80% confidence 80 percent, High |
| 143 | MiniMax-M3 minimax | 81.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.3 | 100% confidence 100 percent, Full |
| 144 | Qwen3.5 35B-A3B Alibaba / Qwen | 81.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.2 | 64% confidence 64 percent, Medium |
| 145 | Qwen3.7 Flash Alibaba / Qwen | 80.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.8 | 32% confidence 32 percent, Low |
| 146 | o3 OpenAI | 80.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.8 | 87% confidence 87 percent, High |
| 147 | Claude Opus 4.5 Anthropic | 80.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.7 | 100% confidence 100 percent, Full |
| 148 | Claude Sonnet 5 Anthropic | 80.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.3 | 93% confidence 93 percent, High |
| 149 | o3 OpenAI | 79.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 79.8 | 87% confidence 87 percent, High |
| 150 | o4-mini OpenAI | 79.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 79.6 | 87% confidence 87 percent, High |
| 151 | Qwen3.5 9B Alibaba / Qwen | 79.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 79.0 | 32% confidence 32 percent, Low |
| 152 | Qwen3.5 9B Alibaba / Qwen | 78.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.9 | 32% confidence 32 percent, Low |
| 153 | Claude Fable 5 Anthropic | 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.8 | 100% confidence 100 percent, Full |
| 154 | Claude Sonnet 4.5 Anthropic | 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Oct 28, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.8 | 80% confidence 80 percent, High |
| 155 | Claude Sonnet 4.6 Anthropic | 78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.8 | 100% confidence 100 percent, Full |
| 156 | Claude Sonnet 3.7 Anthropic | 78.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished May 26, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.5 | 100% confidence 100 percent, Full |
| 157 | GPT-5.4 nano OpenAI | 78.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.5 | 100% confidence 100 percent, Full |
| 158 | Claude Sonnet 4 Anthropic | 78.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 159 | Claude Sonnet 4 Anthropic | 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished May 26, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.8 | 100% confidence 100 percent, Full |
| 160 | o4-mini OpenAI | 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.8 | 87% confidence 87 percent, High |
| 161 | Claude Opus 4.1 Anthropic | 77.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Aug 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.3 | 80% confidence 80 percent, High |
| 162 | GPT-5.5 OpenAI | 77.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.3 | 100% confidence 100 percent, Full |
| 163 | GPT-5.6 Terra OpenAI | 77.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.3 | 100% confidence 100 percent, Full |
| 164 | Claude Sonnet 3.7 Anthropic | 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 76.8 | 100% confidence 100 percent, Full |
| 165 | Claude Sonnet 3.7 Anthropic | 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 10, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 76.8 | 100% confidence 100 percent, Full |
| 166 | Claude Opus 4.1 Anthropic | 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 27KPublished Aug 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 76.8 | 80% confidence 80 percent, High |
| 167 | DeepSeek-R1 DeepSeek | 76.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 76.3 | 88% confidence 88 percent, High |
| 168 | Claude Opus 4 Anthropic | 76.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 76.3 | 100% confidence 100 percent, Full |
| 169 | Claude Sonnet 4 Anthropic | 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.8 | 100% confidence 100 percent, Full |
| 170 | Gemini 3.5 Flash Lite Google | 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.8 | 100% confidence 100 percent, Full |
| 171 | Gemma 4 31B IT Google | 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.8 | 64% confidence 64 percent, Medium |
| 172 | GPT OSS 120B OpenAI | 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.8 | 79% confidence 79 percent, Medium |
| 173 | Nemotron 3.5 Lightning 30B A3B nvidia | 75.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; no tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 75.4 | 100% confidence 100 percent, Full |
| 174 | o4-mini OpenAI | 75.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.3 | 87% confidence 87 percent, High |
| 175 | GPT-5 Mini OpenAI | 75%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.0 | 99% confidence 99 percent, High |
| 176 | GPT-5.4 OpenAI | 74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.7 | 100% confidence 100 percent, Full |
| 177 | Gemini 3.1 Flash Lite Google | 74.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.2 | 32% confidence 32 percent, Low |
| 178 | Gemini 3.5 Flash Lite Google | 74.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.2 | 100% confidence 100 percent, Full |
| 179 | Claude Sonnet 4.5 Anthropic | 73.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Sep 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.7 | 80% confidence 80 percent, High |
| 180 | Gemini 3.1 Flash Lite Google | 73.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.7 | 32% confidence 32 percent, Low |
| 181 | Claude Opus 4.1 Anthropic | 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.2 | 80% confidence 80 percent, High |
| 182 | DeepSeek V4 Pro DeepSeek | 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.2 | 85% confidence 85 percent, High |
| 183 | Gemma 4 26B A4B IT Google | 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.2 | 64% confidence 64 percent, Medium |
| 184 | GPT-5.2 OpenAI | 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.2 | 100% confidence 100 percent, Full |
| 185 | Qwen3 Max Alibaba / Qwen | 72.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 6, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 72.6 | 69% confidence 69 percent, Medium |
| 186 | GPT-5.4 nano OpenAI | 72.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 72.2 | 100% confidence 100 percent, Full |
| 187 | GPT-5 OpenAI | 71.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.7 | 100% confidence 100 percent, Full |
| 188 | GPT-5 Mini OpenAI | 71.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.7 | 99% confidence 99 percent, High |
| 189 | GPT-5 Mini OpenAI | 71.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.7 | 99% confidence 99 percent, High |
| 190 | Claude Haiku 4.5 Anthropic | 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.2 | 64% confidence 64 percent, Medium |
| 191 | DeepSeek Chat DeepSeek | 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.2 | 32% confidence 32 percent, Low |
| 192 | GLM-5.2 Z.ai | 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.2 | 100% confidence 100 percent, Full |
| 193 | Qwen3 235B-A22B Alibaba / Qwen | 70.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 3, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 70.7 | 76% confidence 76 percent, Medium |
| 194 | GPT-5 Nano OpenAI | 69.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 69.4 | 85% confidence 85 percent, High |
| 195 | Claude Opus 4 Anthropic | 69.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 69.2 | 100% confidence 100 percent, Full |
| 196 | DeepSeek V3 0324 DeepSeek | 67.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 1, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 67.6 | 74% confidence 74 percent, Medium |
| 197 | GPT-5 Nano OpenAI | 67.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 67.4 | 85% confidence 85 percent, High |
| 198 | Llama 4 Maverick 17B Instruct Meta | 67.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 67.0 | 78% confidence 78 percent, Medium |
| 199 | GPT-4.1 OpenAI | 66.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.9 | 100% confidence 100 percent, Full |
| 200 | Claude Sonnet 4 Anthropic | 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.7 | 100% confidence 100 percent, Full |
| 201 | GPT-5.1 OpenAI | 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.7 | 85% confidence 85 percent, High |
| 202 | Claude Sonnet 3.7 Anthropic | 66.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.0 | 100% confidence 100 percent, Full |
| 203 | GPT-4.1 mini OpenAI | 65.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.8 | 87% confidence 87 percent, High |
| 204 | Qwen3 32B Alibaba / Qwen | 65.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.7 | 67% confidence 67 percent, Medium |
| 205 | QwQ Plus Alibaba / Qwen | 65.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.4 | 32% confidence 32 percent, Low |
| 206 | QwQ 32B Alibaba / Qwen | 65.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.3 | 67% confidence 67 percent, Medium |
| 207 | DeepSeek-R1-Distill-Qwen-32B DeepSeek | 64.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.1 | 32% confidence 32 percent, Low |
| 208 | GPT-5.4 mini OpenAI | 64.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.1 | 100% confidence 100 percent, Full |
| 209 | GPT-5.6 Luna OpenAI | 63.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 63.6 | 100% confidence 100 percent, Full |
| 210 | Qwen3 30B A3B Alibaba / Qwen | 61.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 61.7 | 64% confidence 64 percent, Medium |
| 211 | GPT OSS 20B OpenAI | 60.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 60.8 | 64% confidence 64 percent, Medium |
| 212 | GLM-4.7-Flash Z.ai | 60.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 60.5 | 69% confidence 69 percent, Medium |
| 213 | Claude Haiku 4.5 Anthropic | 60.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 60.5 | 64% confidence 64 percent, Medium |
| 214 | Mistral Medium 3 mistral | 59.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 59.5 | 85% confidence 85 percent, High |
| 215 | GPT-5 Nano OpenAI | 57.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.6 | 85% confidence 85 percent, High |
| 216 | DeepSeek-V3 DeepSeek | 56.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 56.5 | 93% confidence 93 percent, High |
| 217 | Magistral Small mistral | 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 56.1 | 48% confidence 48 percent, Low |
| 218 | GPT-5.4 nano OpenAI | 55.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.6 | 100% confidence 100 percent, Full |
| 219 | Claude Sonnet 3.5 v2 Anthropic | 55.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.3 | 100% confidence 100 percent, Full |
| 220 | Qwen3 32B Alibaba / Qwen | 54.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 54.1 | 67% confidence 67 percent, Medium |
| 221 | GPT OSS 20B OpenAI | 53.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 53.2 | 64% confidence 64 percent, Medium |
| 222 | Llama 4 Scout 17B Instruct Meta | 51.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 51.8 | 64% confidence 64 percent, Medium |
| 223 | Mistral Large 2.1 mistral | 51.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 51.3 | 100% confidence 100 percent, Full |
| 224 | Qwen3 30B A3B Alibaba / Qwen | 50.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 50.4 | 64% confidence 64 percent, Medium |
| 225 | GPT-4o (2024-08-06) OpenAI | 49.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 49.2 | 100% confidence 100 percent, Full |
| 226 | GPT-4.1 nano OpenAI | 48.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 48.9 | 82% confidence 82 percent, High |
| 227 | GPT-4o (2024-05-13) OpenAI | 48.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 48.9 | 93% confidence 93 percent, High |
| 228 | GPT-5 Nano OpenAI | 48.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 48.5 | 85% confidence 85 percent, High |
| 229 | GPT-4o (2024-11-20) OpenAI | 47.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 47.9 | 61% confidence 61 percent, Medium |
| 230 | Gemma 3 27B IT Google | 47.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 47.7 | 74% confidence 74 percent, Medium |
| 231 | Magistral Small 1.2 mistral | 47.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 47.6 | 32% confidence 32 percent, Low |
| 232 | Llama-3.3-70B-Instruct Meta | 47.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 47.4 | 100% confidence 100 percent, Full |
| 233 | GPT OSS 20B OpenAI | 46.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.0 | 64% confidence 64 percent, Medium |
| 234 | GLM-4.7-Flash Z.ai | 45.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 45.1 | 69% confidence 69 percent, Medium |
| 235 | Llama-3.1-70B-Instruct Meta | 44.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 44.2 | 93% confidence 93 percent, High |
| 236 | Qwen Turbo Alibaba / Qwen | 41.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 41.8 | 47% confidence 47 percent, Low |
| 237 | Gemma 3 12B IT Google | 39.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 39.5 | 64% confidence 64 percent, Medium |
| 238 | Claude Haiku 3.5 Anthropic | 38.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 12, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 38.1 | 100% confidence 100 percent, Full |
| 239 | GPT-4o mini OpenAI | 37.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 37.7 | 100% confidence 100 percent, Full |
| 240 | Claude Haiku 3 Anthropic | 36.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 36.3 | 93% confidence 93 percent, High |
| 241 | Llama-3.1-8B-Instruct Meta | 27.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 27.0 | 93% confidence 93 percent, High |
| 242 | Llama-3.2-1B Meta | 23.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 23.9 | 93% confidence 93 percent, High |
| 243 | Gemma 3 4B IT Google | 23.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 23.2 | 64% confidence 64 percent, Medium |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.