Track one capability without combining different benchmarks into a single series.
Select an evaluation and move through reported results by date. Every point keeps its model, conditions and source. When the protocol changes, the visualization marks the discontinuity.
Loading…
—
How to read the series
—
The slope summarizes published outcomes. It does not isolate model improvement from prompt changes, reasoning effort, harness changes or evaluation conditions.
Active sources
| Date | Model | Result | Protocol | Conditions | Source |
|---|
Each series preserves the protocol, conditions and source for every result.
GPQA Diamond, MMMU-Pro, SWE-bench Verified, SWE-Bench Pro, Toolathlon and ARC-AGI-2 remain separate. They are not converted into a common index.
When a source changes problem count, harness or benchmark version, the series preserves that context. SWE-bench Verified shows two different GPT-5 results under different protocols.
Every observation links to the release table that published the value and keeps the reported conditions. The dataset is reviewed on relevant releases and at least every 30 days while active.
A benchmark curve describes evidence; it does not define an LLM by itself.
To interpret these series alongside architecture, tokens, training and inference, see what an LLM is and how it works. Keeping the two levels separate avoids turning one benchmark into a general definition of model capability.
Comparisons are included only when the evaluation conditions are documented.
This version uses OpenAI release tables because they provide multiple generations with documented evaluation conditions. It does not place other vendors on the same line unless the configuration can be justified as comparable. The dataset can add them later as separate series or protocols when the evidence is sufficient. GPT-6 Astra is added only to GPQA Diamond, where OpenAI publishes a directly comparable value. GPT-6 Sol/Luna, GPT-6.1 Sol, Claude Sonnet 5.5 and Gemini 4 Argon are recorded as reviewed but not added when their tables or protocols are not directly comparable with the existing series.