frontiermath tier 4 v2 Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: math · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-6.1 Sol OpenAI 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
100.0 100% confidence 100 percent, Full
2 GPT-6 Astra OpenAI 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.6 100% confidence 100 percent, Full
3 GPT-6 Astra OpenAI 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.6 100% confidence 100 percent, Full
4 GPT-6 Astra OpenAI 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.6 100% confidence 100 percent, Full
5 GPT-6 Astra OpenAI 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.6 100% confidence 100 percent, Full
6 Claude Opus 5.5 Anthropic 95%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.0 100% confidence 100 percent, Full
7 Claude Fable 5 Anthropic 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.2 100% confidence 100 percent, Full
8 GPT-6 Sol OpenAI 90%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.0 100% confidence 100 percent, Full
9 Claude Fable 5.1 Anthropic 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 1, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.8 100% confidence 100 percent, Full
10 GPT-6 Astra OpenAI 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.8 100% confidence 100 percent, Full
11 GPT-5.6 Sol OpenAI 82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.9 100% confidence 100 percent, Full
12 GPT-6 Astra OpenAI 82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
82.9 100% confidence 100 percent, Full
13 Claude Sonnet 5.5 Anthropic 80.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
80.5 93% confidence 93 percent, High
14 GPT-5.6 Sol OpenAI 80.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] promaxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
80.5 100% confidence 100 percent, Full
15 GPT-5.5 Pro OpenAI 78.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.0 53% confidence 53 percent, Medium
16 Claude Opus 5 Anthropic 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.2 100% confidence 100 percent, Full
17 GPT-5.5 OpenAI 72.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
72.5 100% confidence 100 percent, Full
18 GPT-5.6 Terra OpenAI 70.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
70.7 100% confidence 100 percent, Full
19 GPT-5.6 Luna OpenAI 61.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
61.0 100% confidence 100 percent, Full
20 GPT-5.4 Pro OpenAI 58.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
58.5 69% confidence 69 percent, Medium
21 Claude Opus 4.8 Anthropic 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.1 100% confidence 100 percent, Full
22 GPT-6 Luna OpenAI 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.1 100% confidence 100 percent, Full
23 GPT-5.4 OpenAI 49%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
49.0 100% confidence 100 percent, Full
24 Qwen3.8 Max Alibaba / Qwen 46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
46.3 85% confidence 85 percent, High
25 Muse Spark 1.3 Meta 46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 18, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
46.3 93% confidence 93 percent, High
26 GPT-5.2 Pro OpenAI 46%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
46.0 48% confidence 48 percent, Low
27 Muse Spark 1.3 Meta 41.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 16, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
41.5 93% confidence 93 percent, High
28 Kimi K3 Moonshot AI 39.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
39.0 100% confidence 100 percent, Full
29 Gemini 3.7 Flash Google 36.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
36.6 100% confidence 100 percent, Full
30 Qwen3.7 Max Alibaba / Qwen 34.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
34.1 53% confidence 53 percent, Medium
31 Qwen3.8 Max 0902 Alibaba / Qwen 34.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
34.1 40% confidence 40 percent, Low
32 Claude Opus 4.7 Anthropic 31.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
31.7 100% confidence 100 percent, Full
33 Grok 4.6 xAI 31.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
31.7 100% confidence 100 percent, Full
34 GPT-5.2 OpenAI 31.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
31.7 100% confidence 100 percent, Full
35 Claude Sonnet 5 Anthropic 29.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
29.3 93% confidence 93 percent, High
36 GLM-5.2 Z.ai 29.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 19, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
29.3 100% confidence 100 percent, Full
37 GLM-5.3 Z.ai 29.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 25, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
29.3 93% confidence 93 percent, High
38 Claude Opus 4.6 Anthropic 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
26.8 100% confidence 100 percent, Full
39 DeepSeek V4 Pro 0813 DeepSeek 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 19, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
26.8 69% confidence 69 percent, Medium
40 Gemini 3.1 Pro Preview Google 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
26.8 100% confidence 100 percent, Full
41 Gemini 3.5 Flash Google 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
26.8 100% confidence 100 percent, Full
42 Kimi K2.6 Moonshot AI 25.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
25.6 85% confidence 85 percent, High
43 DeepSeek V4 Flash 0731 DeepSeek 24.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
24.4 69% confidence 69 percent, Medium
44 Grok 4.5 xAI 24.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 9, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
24.4 100% confidence 100 percent, Full
45 Gemini 3.6 Flash Google 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
22.0 100% confidence 100 percent, Full
46 Gemini 3.8 Flash Google 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
22.0 100% confidence 100 percent, Full
47 GPT-5 OpenAI 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
22.0 100% confidence 100 percent, Full
48 GPT-5 Pro OpenAI 19.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
19.5 64% confidence 64 percent, Medium
49 Gemini 3 Flash Preview Google 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
17.1 93% confidence 93 percent, High
50 Inkling Small thinkingmachines 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
17.1 100% confidence 100 percent, Full
51 Grok 4.20 (Reasoning) xAI 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
17.1 80% confidence 80 percent, High
52 Grok 4.7 xAI 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 22, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
17.1 100% confidence 100 percent, Full
53 GLM-5.3-Flash Z.ai 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 27, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
17.1 100% confidence 100 percent, Full
54 Grok 4.3 xAI 14.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
14.6 80% confidence 80 percent, High
55 Kimi K2.7 Code Moonshot AI 12.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 13, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
12.2 48% confidence 48 percent, Low
56 GPT-5 Mini OpenAI 12.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
12.2 99% confidence 99 percent, High
57 GPT-5.4 nano OpenAI 12.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
12.2 100% confidence 100 percent, Full
58 GPT-5.4 mini OpenAI 9.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
9.8 100% confidence 100 percent, Full
59 Claude Opus 4.5 Anthropic 4.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
4.9 100% confidence 100 percent, Full
60 o4-mini OpenAI 4.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
4.9 87% confidence 87 percent, High
61 Inkling thinkingmachines 4.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
4.9 100% confidence 100 percent, Full
62 Claude Opus 4.1 Anthropic 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
2.4 80% confidence 80 percent, High
63 Claude Sonnet 4.5 Anthropic 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
2.4 80% confidence 80 percent, High
64 DeepSeek V4 Pro DeepSeek 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
2.4 85% confidence 85 percent, High
65 GPT-5 Nano OpenAI 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
2.4 85% confidence 85 percent, High
66 GPT-5.5 Instant OpenAI 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
2.4 64% confidence 64 percent, Medium
67 Gemini 2.5 Pro Google 0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
0.0 90% confidence 90 percent, High
68 Gemini 3.5 Flash Lite Google 0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
0.0 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the math pillar of the SI Score.