The harder the benchmark, the more the frontier keeps
Best score from any lab outside Anthropic and OpenAI, minus the best Anthropic/OpenAI score, one dot per benchmark, against the benchmark's fitted difficulty. Each week is the state of the store as models became available. Third-party scores only.
Source: Synopticon score store, 42,156 records across 49 sources · IRT fit, 950 models · 2026-08-15 · live version