Top Super Intelligence Models: 10 Frontier AI Profiles

A shortlist of the ten highest SI Scores, with measured strengths, price, context, confidence and open-weight status in one place.

Snapshot: Oct 9, 2026, 01:45 UTC · Method si-v3-retained-evidence-2

A shortlist you can actually compare

These profiles follow the current overall ranks, rather than treating a famous brand or a new release as evidence of superiority. Each card identifies the strongest measured pillars within that model, lists the catalog facts available today, and links to the underlying results. “Strongest” compares the model’s own pillar scores; it does not mean that it wins that pillar against every competitor. The SI Score beside each profile is adjusted for evidence support.

The top ten are rank-eligible models, not every model with a score. A provisional score remains outside this shortlist until the evidence meets the published ranking rule. There may be strong models outside the ten, including specialized systems with only one reported capability. Treat these cards as a broad frontier comparison and use task rankings when you have a specific job to accomplish.

#1 · Anthropic

Claude Fable 5.1

80.2SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (91.5) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$10.00 / $50.00
Context window
1M
Weights
Closed weights
Evidence
50 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#2 · Anthropic

Claude Opus 5.5

78.4SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (91.5) and reasoning (91.4). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$4.00 / $20.00
Context window
1M
Weights
Closed weights
Evidence
45 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#3 · OpenAI

GPT-6 Astra

77.6SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (92.9) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$10.00 / $50.00
Context window
1.1M
Weights
Closed weights
Evidence
63 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#4 · Anthropic

Claude Fable 5

76.8SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (91.0) and reasoning (89.4). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$10.00 / $50.00
Context window
1M
Weights
Closed weights
Evidence
51 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#5 · Anthropic

Claude Opus 5

74.9SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (86.9) and reasoning (84.3). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$5.00 / $25.00
Context window
1M
Weights
Closed weights
Evidence
46 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#6 · OpenAI

GPT-6.1 Sol

73.0SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (93.0) and reasoning (87.5). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$2.00 / $10.00
Context window
1.1M
Weights
Closed weights
Evidence
47 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#7 · OpenAI

GPT-5.6 Sol

72.1SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (89.3) and reasoning (82.1). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$4.00 / $20.00
Context window
1.1M
Weights
Closed weights
Evidence
56 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#8 · Anthropic

Claude Opus 4.7

71.2SI Score · 100% confidence 100 percent, Full

Highest measured pillars: preference (80.5) and math (78.0). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$5.00 / $25.00
Context window
1M
Weights
Closed weights
Evidence
47 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#9 · Moonshot AI

Kimi K3

71.0SI Score · 100% confidence 100 percent, Full

Highest measured pillars: reasoning (81.3) and preference (79.9). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$3.00 / $15.00
Context window
1M
Weights
Open weights
Evidence
38 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

#10 · Anthropic

Claude Opus 4.6

70.8SI Score · 100% confidence 100 percent, Full

Highest measured pillars: preference (81.6) and reasoning (78.8). These describe published evaluations, rather than a guarantee on your workload.

Input / output per 1M tokens
$5.00 / $25.00
Context window
1M
Weights
Closed weights
Evidence
46 published result rows

Prices and context are catalog or provider facts, not measurements by SI Index. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.

Inspect benchmarks and sources →

Read price and context together

API prices are per million tokens, split into input and output. An inexpensive input rate can coexist with an expensive output rate. Long conversations, retrieved documents, tool outputs and generated code change the balance. Compare both rates, then calculate a bill using your expected token mix. Prices do not include your application’s infrastructure, and a missing figure should be treated as an unresolved cost rather than a free service.

The context window is an advertised limit, not a test of reliable recall across that entire window. A large window can accommodate more material but does not prove that a model will find the relevant detail or follow every constraint. For document-heavy workflows, test the actual document length and placement of critical information. The long-context board orders the known limits; the linked source establishes what each limit means.

Evidence before commitment

Downloadable weights can make local hosting or customization possible, while licenses and hardware determine whether those options fit your use case. Confidence tells you how much expected benchmark-source weight has arrived, not whether a model is safe or suitable for a particular deployment. Open each profile before choosing: result notes retain the benchmark variant, the source and the retrieval date. The comparison pages put shared tests beside each other so a difference is easier to inspect.

Explore SI Index