terminal bench v4 0 Benchmark: Scores and Sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: coding · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Claude Sonnet 5.5 Anthropic 70.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
70.6 93% confidence 93 percent, High
2 Claude Opus 5.5 Anthropic 66.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; production safeguards with fallback; 4.0Published Sep 22, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
66.4 100% confidence 100 percent, Full
3 Claude Opus 5.5 Anthropic 64.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
64.8 100% confidence 100 percent, Full
4 Claude Sonnet 5.5 Anthropic 61.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
61.8 93% confidence 93 percent, High
5 GPT-6 Astra OpenAI 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 10, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
58.2 100% confidence 100 percent, Full
6 GPT-6.1 Sol OpenAI 58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
58.2 100% confidence 100 percent, Full
7 GPT-6 Astra OpenAI 57.9%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
57.9 100% confidence 100 percent, Full
8 Claude Fable 5.1 Anthropic 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
57.9 100% confidence 100 percent, Full
9 Claude Fable 5.1 Anthropic 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhighPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
57.9 100% confidence 100 percent, Full
10 GPT-6 Astra OpenAI 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; highPublished Sep 10, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
57.9 100% confidence 100 percent, Full
11 GPT-6 Astra OpenAI 57.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; xhighPublished Sep 10, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
57.9 100% confidence 100 percent, Full
12 Claude Fable 5.1 Anthropic 55.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] production safeguards with fallback; 4.0Published Sep 1, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
55.8 100% confidence 100 percent, Full
13 Claude Fable 5.1 Anthropic 54.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; highPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
54.5 100% confidence 100 percent, Full
14 GPT-6 Astra OpenAI 54.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; mediumPublished Sep 10, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
54.2 100% confidence 100 percent, Full
15 Claude Fable 5.1 Anthropic 53.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; mediumPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
53.9 100% confidence 100 percent, Full
16 Claude Opus 5 Anthropic 53.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhighPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
53.9 100% confidence 100 percent, Full
17 Claude Opus 5 Anthropic 51.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
51.8 100% confidence 100 percent, Full
18 GPT-6 Astra OpenAI 50.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; lowPublished Sep 10, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
50.6 100% confidence 100 percent, Full
19 Claude Opus 5 Anthropic 50.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; highPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
50.3 100% confidence 100 percent, Full
20 GPT-6 Sol OpenAI 49.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
49.4 100% confidence 100 percent, Full
21 Claude Opus 5 Anthropic 44.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; mediumPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
44.9 100% confidence 100 percent, Full
22 Claude Fable 5 Anthropic 44.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
44.5 100% confidence 100 percent, Full
23 Claude Fable 5.1 Anthropic 43.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; lowPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
43.3 100% confidence 100 percent, Full
24 GLM-5.3 Z.ai 41.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
41.8 93% confidence 93 percent, High
25 Claude Haiku 5.5 Anthropic 39.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
39.2 100% confidence 100 percent, Full
26 Grok 4.7 xAI 37.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; 4.0Published Sep 21, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
37.6 100% confidence 100 percent, Full
27 Grok 4.7 xAI 37.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Grok Build; xhighPublished Sep 21, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
37.6 100% confidence 100 percent, Full
28 GPT-5.6 Sol OpenAI 37.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
37.3 100% confidence 100 percent, Full
29 GLM-5.3-Flash Z.ai 35.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
35.8 100% confidence 100 percent, Full
30 MiMo-V2.6-Pro xiaomi 34.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.6 Pro comparison column; 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
34.9 88% confidence 88 percent, High
31 Claude Opus 5 Anthropic 34.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; lowPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
34.9 100% confidence 100 percent, Full
32 DeepSeek V4.1 Flash DeepSeek 31.2%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness minimal mode; 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
31.2 69% confidence 69 percent, Medium
33 MiMo-V2.6-Flash xiaomi 28.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
28.8 88% confidence 88 percent, High
34 Qwen3.8 Max 0902 Alibaba / Qwen 27.0%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
27.0 40% confidence 40 percent, Low
35 Claude Opus 4.8 Anthropic 23.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
23.6 100% confidence 100 percent, Full
36 GPT-5.6 Terra OpenAI 21.5%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
21.5 100% confidence 100 percent, Full
37 Grok 4.6 xAI 20.3%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] high effort; 4.0Published Sep 21, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
20.3 100% confidence 100 percent, Full
38 Grok 4.6 xAI 20.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Grok Build; highPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
20.3 100% confidence 100 percent, Full
39 Gemini 3.8 Flash Google 19.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
19.1 100% confidence 100 percent, Full
40 Gemini 3.8 Flash Google 19.1%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] mini-SWE-agent; highPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
19.1 100% confidence 100 percent, Full
41 GPT-5.6 Luna OpenAI 17.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
17.3 100% confidence 100 percent, Full
42 GPT-6 Luna OpenAI 16.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
16.4 100% confidence 100 percent, Full
43 Muse Spark 1.3 Meta 14.6%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Muse Code; xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
14.6 93% confidence 93 percent, High
44 Claude Sonnet 5 Anthropic 12.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
12.4 93% confidence 93 percent, High
45 Grok 4.5 xAI 12.4%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Grok Build; highPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
12.4 100% confidence 100 percent, Full
46 Gemini 3.7 Flash Google 11.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] mini-SWE-agent; highPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
Open source ↗
11.2 100% confidence 100 percent, Full
47 MiMo-V2.5-Pro xiaomi 1.5%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.5 Pro comparison column; 4.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
1.5 88% confidence 88 percent, High

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.