The SI Score, in plain words
The SI Score is our composite ranking of Super Intelligence models, on a 0–100 scale. We do not run our own evaluations. Instead, we collect 57 published benchmarks from open sources, group them into 4 capability pillars, and combine them with published weights:
| Pillar | Benchmarks | Weight in SI Score |
|---|---|---|
| reasoning | arc agi 1 (1), arc agi 2 (1), arc agi 3 (1), arc agi v1 public eval (1), arc agi v1 semi private (1), arc agi v2 public eval (1), arc agi v2 semi private (1), arc agi v3 semi private (1), gpqa diamond (0.5), hle (1), hle 1811 verified and revised items (1), hle full set (1), hle full set text mm (1), hle full set tools (1), hle scale (1), hle text only (1), hle text only subset (1), hle text only subset tools (1), hle text only tools (1), hle tools (1), livebench reasoning connections 2026_06_25 (1), livebench reasoning consecutive events 2026_06_25 (1), livebench reasoning logic with navigation 2026_06_25 (1), livebench reasoning spatial 2026_06_25 (1), livebench reasoning theory of mind 2026_06_25 (1), livebench reasoning zebra puzzle 2026_06_25 (1), mmlu pro (0.5) | 30% |
| math | frontiermath tier 1 3 (1), frontiermath tier 4 (1), frontiermath tier 4 v2 (1), frontiermath tiers 1 3 v2 (1), frontiermath vv2 tier 1 3 (1), frontiermath vv2 tier 4 (1), livebench math amps hard 2026_06_25 (1), livebench math integrals with game 2026_06_25 (1), livebench math math comp 2026_06_25 (1), livebench math olympiad 2026_06_25 (1), livebench math simplify 2026_06_25 (1), math level 5 (1), otis mock aime 2024 2025 (0.5) | 15% |
| coding | aider polyglot (1), livebench coding code completion 2026_06_25 (1), livebench coding code generation 2026_06_25 (1), livebench coding javascript 2026_06_25 (1), livebench coding python 2026_06_25 (1), livebench coding typescript 2026_06_25 (1), swe bench pro (1), swe bench pro public (1), swe bench pro qwen refined and corrected task set (1), swe bench verified (1), terminal bench (1), terminal bench v0 1 (1), terminal bench v2 0 (1), terminal bench v2 1 (1), terminal bench v3 0 (1), terminal bench v4 0 (1) | 40% |
| preference | lmarena text (1) | 15% |
The production method si-v2-absolute-shrinkage-1 uses fixed absolute scales, not percentiles.
Percentage results retain their 0–100 value (inverted for lower-is-better metrics). Elo is mapped
through a logistic curve centered at 1,200 with scale 200; time values use 100 × seconds ÷
(seconds + 30), with direction applied afterward. Adding a model does not change another model's
normalized result. Versions, tools, effort and harness conditions remain separate benchmark IDs or result notes.
Multiple submissions from one source for a benchmark are averaged first; source means are combined with reliability weights. Lab-reported evidence has weight 0.35; inactive sources receive a 0.5 factor. Each pillar averages its reported benchmarks using benchmark weight and evidence reliability. Missing pillars stay unknown; this method does not impute them.
The reported pillar mean is reweighted over available pillars, then shrunk toward a neutral prior of 50. Support is the minimum of expected-source coverage, reported pillar weight, and evidence breadth (weighted benchmark evidence divided by a target of four, capped at one). SI Score = 50 + support × (reported pillar mean − 50). This reduces the influence of sparse evidence; it does not remove all differences in evaluation conditions. A rank requires at least 50% confidence and two scored pillars. The fixture used during development describes an older method; production scores always come from the pipeline.
The confidence %, in plain words
Sources publish at different times: an arena board may list a new model on day one, a benchmark steward may take a week, and some sources never cover some models. So every score on this site carries a confidence percentage that starts low and rises as expected sources report.
Each model has a set of expected sources with weights (for example, a frontier model is expected on active benchmark sources whose publications postdate its release; catalog and pricing sources do not contribute to capability confidence). The model's coverage is the share of expected weight that has arrived. Confidence is coverage divided by a completeness threshold of 80%, capped at 100:
confidence = min(100, 100 × coverage ÷ 0.8)
In words: a model shows 100% confidence once about 80% of its expected source weight has reported — we do not wait for the last stragglers, because some sources never cover some models. Example: Qwen2.5-Coder-32B-Instruct has pending sources: Official model cards via models.dev and LiveBench; its coverage is 60%, so it shows 75% confidence. Model pages list exactly which sources are in (with arrival dates) and which are pending, and the status page lists every model waiting on any source.
Sources and licenses
Every displayed value links to the source it came from, with the retrieval date. 5 of our 22 sources publish under an open license; the rest are cited as factual figures (prices, release dates, benchmark results) from the publisher's own page.
| Source | Kind | License | Attribution |
|---|---|---|---|
| models.dev | catalog | MIT | models.dev; MIT; source links accompany every observation. |
| Official model cards via models.dev | lab-reported | factual citation; MIT transcription | Official model cards via models.dev; factual citation; MIT transcription; source links accompany every observation. |
| LiteLLM | pricing | MIT | LiteLLM; MIT; source links accompany every observation. |
| Epoch AI Benchmarking | benchmark | CC-BY | Epoch AI Benchmarking; CC-BY; source links accompany every observation. |
| LMArena / Arena | preference | CC-BY-4.0 | LMArena / Arena; CC-BY-4.0; source links accompany every observation. |
| SWE-bench Verified | benchmark | factual citation | SWE-bench Verified; factual citation; source links accompany every observation. |
| SWE-bench Pro (public) | benchmark | factual citation | SWE-bench Pro (public); factual citation; source links accompany every observation. |
| ARC Prize | benchmark | factual citation | ARC Prize; factual citation; source links accompany every observation. |
| Humanity’s Last Exam | benchmark | factual citation | Humanity’s Last Exam; factual citation; source links accompany every observation. |
| LiveBench | benchmark | factual citation; Apache-2.0 code | LiveBench; factual citation; Apache-2.0 code; source links accompany every observation. |
| Aider polyglot | benchmark | Apache-2.0 | Aider polyglot; Apache-2.0; source links accompany every observation. |
| Terminal-Bench | benchmark | Apache-2.0; factual citation | Terminal-Bench; Apache-2.0; factual citation; source links accompany every observation. |
| OpenRouter rankings | popularity | CC-BY-4.0 | OpenRouter rankings; CC-BY-4.0; source links accompany every observation. |
| Hugging Face Hub | catalog | factual metadata; model-specific licenses | Hugging Face Hub; factual metadata; model-specific licenses; source links accompany every observation. |
| Anthropic models & pricing | pricing | factual citation | Anthropic models & pricing; factual citation; source links accompany every observation. |
| OpenAI API changelog | catalog | factual citation | OpenAI API changelog; factual citation; source links accompany every observation. |
| Google Gemini release notes | catalog | CC-BY-4.0 factual citation | Google Gemini release notes; CC-BY-4.0 factual citation; source links accompany every observation. |
| DeepSeek V4.1 Flash announcement | catalog | factual citation | DeepSeek V4.1 Flash announcement; factual citation; source links accompany every observation. |
| xAI models & pricing | pricing | factual citation | xAI models & pricing; factual citation; source links accompany every observation. |
| Provider pricing pages and model cards | pricing / lab-reported | factual citation | linked individually next to each value |
What we deliberately do not use: Artificial Analysis and llm-stats are consulted only as an internal cross-check of our own normalized values — their terms restrict reuse in a competing product, so none of their numbers are published here. Benchmark datasets and questions themselves are third-party IP and are never republished; only results are cited.
Update policy
New models are detected from catalog diffs and provider announcements, and go live as soon as identity and a primary source are confirmed — marked provisional, with unknown fields shown as “not yet reported” rather than borrowed from a sibling model. Scores update as sources report, and the “as of” date in the header shows when the current snapshot was generated. Prices retain the linked catalog or provider provenance; missing prices are not estimated.
Corrections
Every figure keeps its source and retrieval date, so errors are traceable. When a source corrects a figure we update it and keep the change in the snapshot history; this launch snapshot does not yet publish per-model change histories. Corrections should identify the model, benchmark variant, source URL and retrieval date so they can be reproduced.
Reusing our data
Our composite — the SI Score, confidence and ranks — is published under CC-BY-4.0 with attribution to SI Index — the Superintelligence Leaderboard:
- /data/latest.json — the full composite, machine-readable
- /data/models.csv — composite + key facts per model
- /data/benchmarks.csv — the open-licensed benchmark inputs, with per-row source, date and license
Values from “factual citation” sources (for example SWE-bench or provider pricing pages) are shown on this site with attribution but are not included in the open-inputs CSV, because we do not hold a redistribution license for them. Check the original publisher's terms before reusing those.