GPT-6 Astra vs GPT-6.1 Sol
VS
SI Score, rank and confidence side by side
Pillars — GPT-6 Astra
Coding (weight 40 percent) 64.8 / 65.0
Math (weight 15 percent) 92.9 / 93.0
Preference (weight 15 percent) 77.1 / 77.3
Reasoning (weight 30 percent) 87.1 / 87.5
Right-hand value is GPT-6.1 Sol.
The basics
| Attribute | GPT-6 Astra | GPT-6.1 Sol |
|---|---|---|
| Input price / 1M | $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $2.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Output price / 1M | $50.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Context window | 1.1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ | 1.1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Released | Sep 3, 2026OpenAI API changelogPublished source fact
Retrieved Oct 9, 2026 · factual citation Open source ↗ | Sep 29, 2026OpenAI API changelogPublished source fact
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Open weights | Closed | Closed |
Shared benchmarks
27 in common| Benchmark | GPT-6 Astra | GPT-6.1 Sol | Conditions |
|---|---|---|---|
| arc agi v1 public eval | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-high 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-low 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-max 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-medium 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhigh | Check variant and harness |
| arc agi v1 semi private | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-high 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-low 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-max 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-medium 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhigh | Check variant and harness |
| arc agi v2 public eval | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 96.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-high 77.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-low 97.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-max 90.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-medium 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhigh | Check variant and harness |
| arc agi v2 semi private | 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-high 76.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-low 94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-max 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-medium 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhigh | Check variant and harness |
| arc agi v3 semi private | 54.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-high 17.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-low 62.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-max 38.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-medium 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhigh | 26.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-high 3.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-low 52.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-max 10.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-medium 39.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ ARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhigh | Check variant and harness |
| frontiermath tier 4 v2 | 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] high 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] low 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] medium 82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] none 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhigh | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| frontiermath tiers 1 3 v2 | 93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| gpqa diamond | 95.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max 96%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] | 95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| livebench coding code completion 2026_06_25 | 80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding code generation 2026_06_25 | 80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 78.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding javascript 2026_06_25 | 63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding python 2026_06_25 | 65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench coding typescript 2026_06_25 | 43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 46.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math amps hard 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math integrals with game 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math math comp 2026_06_25 | 97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math olympiad 2026_06_25 | 92.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 91.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench math simplify 2026_06_25 | 70.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning connections 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning consecutive events 2026_06_25 | 90.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 91.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning logic with navigation 2026_06_25 | 88%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 82%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning spatial 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning theory of mind 2026_06_25 | 84.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 86.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| livebench reasoning zebra puzzle 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code Open source ↗ LiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25 | Check variant and harness |
| lmarena text | 1442.3 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 2, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0 Open source ↗ Published source fact [variant] text / overall | 1445.6 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 2, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0 Open source ↗ Published source fact [variant] text / overall | Check variant and harness |
| otis mock aime 2024 2025 | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ Epoch-owned evaluation mean score; scores only, no benchmark questions [variant] max | Check variant and harness |
| terminal bench v4 0 | 57.9%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ Lab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; highPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; high 50.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; lowPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; low 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; max 54.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; mediumPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; medium 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhighPublished Sep 10, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhigh | 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · Apache-2.0; factual citation Open source ↗ Terminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; max | Check variant and harness |
Only benchmarks both models have are compared. Raw values are in each benchmark's own unit; a direct comparison requires matching evaluation conditions. All reported source/variant rows are shown; no raw-result winner is assigned across unmatched harnesses. Sources and dates sit behind every dotted number.
Which should you choose?
What the data says — heuristics from the numbers above, not a verdict:
- GPT-6.1 Sol has the lower input-token price ($2.00 vs $10.00 per 1M).
- Confidence differs: GPT-6 Astra 100% vs GPT-6.1 Sol 100% — pending sources can still move either score.