Best AI Models for Math
Mathematical-reasoning models ordered by the math pillar from the available published evaluations.
Current inputs: frontiermath tier 1 3 (weight 1); frontiermath tier 4 (weight 1); frontiermath tier 4 v2 (weight 1); frontiermath tiers 1 3 v2 (weight 1); frontiermath vv2 tier 1 3 (weight 1); frontiermath vv2 tier 4 (weight 1); livebench math amps hard 2026_06_25 (weight 1); livebench math integrals with game 2026_06_25 (weight 1); livebench math math comp 2026_06_25 (weight 1); livebench math olympiad 2026_06_25 (weight 1); livebench math simplify 2026_06_25 (weight 1); math level 5 (weight 1); otis mock aime 2024 2025 (weight 0.5). Pillar scores can have different evidence coverage; review the individual results.
| # | Model | Math pillar | SI Score | Confidence |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol OpenAI | 93.0 | 73.0 | 100% confidence 100 percent, Full |
| 2 | GPT-6 Astra OpenAI | 92.9 | 77.6 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5.5 Anthropic | 91.5 | 78.4 | 100% confidence 100 percent, Full |
| 4 | Claude Fable 5.1 Anthropic | 91.5 | 80.2 | 100% confidence 100 percent, Full |
| 5 | Gemini 3 Pro Preview Google provisional | 91.4 | 58.6 | 68% confidence 68 percent, Medium |
| 6 | GPT-6 Sol OpenAI | 91.3 | 68.9 | 100% confidence 100 percent, Full |
| 7 | Qwen3.6 Max Preview Alibaba / Qwen provisional | 91.1 | 66.4 | 64% confidence 64 percent, Medium |
| 8 | Claude Fable 5 Anthropic provisional | 91.0 | 76.8 | 100% confidence 100 percent, Full |
| 9 | Claude Sonnet 5.5 Anthropic | 90.5 | 69.4 | 93% confidence 93 percent, High |
| 10 | GPT-5.6 Sol OpenAI provisional | 89.3 | 72.1 | 100% confidence 100 percent, Full |
| 11 | GPT OSS 120B OpenAI provisional | 88.9 | 54.8 | 79% confidence 79 percent, Medium |
| 12 | DeepSeek V4.1 Flash DeepSeek | 88.0 | 66.2 | 69% confidence 69 percent, Medium |
| 13 | Claude Opus 5 Anthropic provisional | 86.9 | 74.9 | 100% confidence 100 percent, Full |
| 14 | Muse Spark 1.2 Meta provisional | 86.7 | 62.3 | 53% confidence 53 percent, Medium |
| 15 | DeepSeek-R1 DeepSeek provisional | 86.6 | 47.9 | 88% confidence 88 percent, High |
Task pages rank on one transparent metric, not the blended SI Score — see methodology for how pillars and confidence are computed. Missing values mean the source has not reported them.