#1 · Anthropic
Claude Fable 5.1
Highest measured pillars: math (91.5) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.
Inspect benchmarks and sources →The current leader, its nearest alternatives, and task-specific choices from the live SI Score snapshot. An evidence-based answer with its limits made explicit.
At the snapshot date above, Claude Fable 5.1 ranks #1 with an SI Score of 80.2 and 100% confidence. “Best superintelligence” here means the highest-ranked frontier AI model under our published method. The ranking does not establish that a beyond-human general intelligence exists.
#1 · Anthropic
Highest measured pillars: math (91.5) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.
Inspect benchmarks and sources →#2 · Anthropic
Highest measured pillars: math (91.5) and reasoning (91.4). These describe published evaluations, rather than a guarantee on your workload.
Inspect benchmarks and sources →#3 · OpenAI
Highest measured pillars: math (92.9) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.
Inspect benchmarks and sources →The leader has the highest composite among models that meet the rank eligibility rule. That rule asks for enough expected-source coverage and results in at least two capability pillars. It prevents an isolated impressive result from receiving a headline rank. The composite combines reasoning, math, coding and human preference using the weights in our methodology. This is our synthesis of published evidence; we do not run the evaluations ourselves.
A first place finish is useful for narrowing a shortlist, but the score gap is not a measured probability that one model will beat another on your next prompt. Different harnesses, reasoning budgets and benchmark versions affect outcomes. Open the three profiles, compare their shared tests, and pay particular attention to the variants that resemble your application. A coding agent operating in a repository and a conversational assistant are different jobs.
Confidence reports expected-source completeness, adjusted by the method’s 80% threshold. It is not a forecast, a safety score or a guarantee of correct answers. A model can reach 100% while some sources remain pending. Missing prices, context or benchmark results stay unknown; they are never borrowed from related versions. A provisional model may be excellent but has insufficient evidence for a comparable overall rank.
For a practical choice, shortlist the leader and a task specialist, test representative inputs, and compare quality with cost and latency in your own environment. SI Index currently does not measure latency. For a plain-language buying decision read best AI model; for the definition behind this query use the FAQ.