Synopticon Capability Index

Every AI model on one ruler.

Benchmark scores don't compare: every test has its own difficulty, its own chance floor, its own grading. The SCI fits them into a single capability score per model, and weighs an independent measurement three times as heavily as the lab's own claim.

The frontier gains about a point a month

Every dated model by release date. The step line tracks the best score to date; the scale is pinned so Claude 3.5 Sonnet = 130 and GPT-5 (high) = 150.

The leaderboard

One row per model; a row folds its reasoning-effort and harness variants together and shows the best configuration. Click a row to see every configuration we score.

# Model Lab SCI Evidence Released

We grade the labs' own numbers, too

Median gap between what a lab reports for its models and what independent evaluators measure on the same benchmarks, in points. Positive means the lab's own numbers run hot.

Like-for-like score pairs only. Labs with at least 6 comparable pairs.
Full benchmark-inflation report →

Capability is getting cheaper

SCI against the cost of running one full evaluation suite (Artificial Analysis, log scale). Down and to the right is the efficient frontier.

Models with a matched Artificial Analysis cost run. Labels mark the efficient frontier.

The harder the benchmark, the more the frontier keeps

One dot per benchmark: the best independent score from any other lab minus the best from any Anthropic or OpenAI model, against the benchmark's fitted difficulty (SDI, same scale as SCI). Above zero the pack leads; below it the frontier labs hold. The line is the median gap by difficulty band.

How the index is built

01 · COLLECT
Every score, with receipts

Daily collection from 40+ sources: public leaderboards, independent evaluators, lab announcements and system cards. Every score keeps its source, date, and eval conditions.

02 · RECONCILE
Same model, same name

The same model ships under dozens of names and configurations. A canonical registry resolves them, so evidence lands on the right model instead of fragmenting.

03 · FIT
One ruler for all of it

A statistical model of test difficulty places every model and every benchmark on one scale: the same approach Epoch AI uses for its capabilities index, fitted over 600+ benchmarks to their ~40.

The fit treats each benchmark like an exam question: hard exams that few models pass say more than saturated ones, and a model's score is the capability level that best explains all of its results at once. Multiple-choice benchmarks are corrected for their guess-rate floor before fitting, so a 25%-chance test and an open-ended one count fairly.

Trust is the design constraint. A lab's own published number enters the fit at one-third the weight of an independent measurement. Models seen only on easy, saturated benchmarks are held out of the ranking rather than flattered by it, and models with fewer than five hard-benchmark results are hidden by default. The benchmark-inflation report is the same machinery pointed at the labs themselves.

The method is Epoch AI's. Their Epoch Capabilities Index (ECI) introduced this fit in A Rosetta Stone for AI Benchmarks, and we anchor to the same two points (Claude 3.5 Sonnet = 130, GPT-5 (high) = 150) so an SCI reads in ECI units. Where the two indices overlap, they agree closely.

What SCI adds is coverage and skepticism: Epoch fits roughly 40 benchmarks; we fit 600+ across 1,200+ tracked models, fold in sources Epoch doesn't ingest, weight independent evidence over self-reported, and rebuild every day. SCI is Synopticon's own independent fit — not Epoch's index, and not affiliated with Epoch AI.

The fine print

What SCI is not. It is not a product review. It measures benchmark capability, not latency, price, context handling, or how a model feels in daily use. A single number also hides real specialization: a model can be frontier at code and mediocre at vision.

Preview models. Rows marked preview are vendor-designated preview or pre-release SKUs. Their scores come from published evaluations of that SKU and may not match the shipping product. A pre-release SKU whose released version is also ranked is folded into that row.

Uncertainty. The daily rebuild publishes point estimates; bootstrapped confidence intervals are computed on periodic manual runs. The evidence count on each row is the honest trust signal; treat thin rows as provisional.