terminal bench v2 0 Benchmark: Scores and Sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-5.5 OpenAI | 82.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.0Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 82.7 | 100% confidence 100 percent, Full |
| 2 | GPT-5.4 OpenAI | 75.1%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.0Published Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 75.1 | 100% confidence 100 percent, Full |
| 3 | Qwen3.7 Max Alibaba / Qwen | 69.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.0Published May 19, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 69.7 | 53% confidence 53 percent, Medium |
| 4 | DeepSeek V4 Pro DeepSeek | 67.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; 2.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 67.9 | 85% confidence 85 percent, High |
| 5 | MiMo-V2.5 xiaomi | 65.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 65.8 | 88% confidence 88 percent, High |
| 6 | GPT-5.4 mini OpenAI | 60%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhigh; 2.0Published Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 60.0 | 100% confidence 100 percent, Full |
| 7 | DeepSeek V4 Flash DeepSeek | 56.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; 2.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 56.9 | 53% confidence 53 percent, Medium |
| 8 | GPT-5.4 nano OpenAI | 46.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhigh; 2.0Published Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 46.3 | 100% confidence 100 percent, Full |
| 9 | Laguna XS 2.1 poolside | 37.5%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Harbor; 2.0Published Jul 2, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 37.5 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.