Benchmark inflation

How hot does each lab run its own scores?

When a lab publishes a benchmark result, is it the number an independent evaluator would get? For every model-and-benchmark pair where we have both a self-reported and a third-party score, we take the gap and report the median per lab — over like-for-like comparisons.

Positive = the lab scores itself higher than independent evaluators do. Negative = the lab is conservative. Only like-for-like pairs count: matched model, matched benchmark, comparable eval conditions. Pairs where the two numbers measure genuinely different things (a different harness, a >15-point methodology gap) are set aside as suspect, not counted as inflation.

Median self-reported minus third-party score, by lab
Points on a 0–100 benchmark scale · like-for-like (gold + silver) pairs only
conservative 0 — matches independents0 scores itself high

The largest credible gaps

Individual like-for-like pairs, sorted by the self-over-third gap. These are the comparisons behind the bars above: same model, same benchmark, a real difference between what the lab reported and what an independent run found.

Model Lab Benchmark Self Third-party Gap Comparability

How the gap is measured

For each model we collapse every self-reported score on a benchmark to one number (the median across sources) and every independent score to another, then take self − third. Before comparing, each pair is graded on how alike the two measurements actually are: gold (gap ≤ 5 points and stable across sources), silver (gap ≤ 15), or bronze (> 15 points, or the eval conditions differ — with-tools versus no-tools, a different harness).

Only gold and silver pairs enter the published gap. A bronze pair — say a 92% self-reported SimpleQA against a 27% third-party number — is a different measurement, not a lie, so counting it as inflation would be misleading. This is what keeps the number defensible: on the raw mean gap one lab looks like a large inflator, but once its non-comparable pairs are set aside its like-for-like gap is near zero, with the divergence flagged separately as suspect.

This is the same score store that feeds the Synopticon Capability Index. There, self-reported scores are down-weighted inside the capability fit; here, the self-versus-independent gap is measured directly as its own published artifact.