Track one capability without combining different benchmarks into a single series.
Select an evaluation and move through reported results by date. Every point keeps its model, conditions and source. When the protocol changes, the visualization marks the discontinuity.
Loading…
—
How to read the series
—
The slope summarizes published outcomes. It does not isolate model improvement from prompt changes, reasoning effort, harness changes or evaluation conditions.
Active sources
| Date | Model | Result | Protocol | Conditions | Source |
|---|
Each series preserves the protocol, conditions and source for every result.
GPQA Diamond, MMMU-Pro, SWE-bench Verified, SWE-Bench Pro, Toolathlon and ARC-AGI-2 remain separate. They are not converted into a common index.
When a source changes problem count, harness or benchmark version, the series preserves that context. SWE-bench Verified shows two different GPT-5 results under different protocols.
Every observation links to the release table that published the value and keeps the reported conditions. The dataset is reviewed on relevant releases and at least every 30 days while active.
Comparisons are included only when the evaluation conditions are documented.
This version uses OpenAI release tables because they provide multiple generations with documented evaluation conditions. It does not place other vendors on the same line unless the configuration can be justified as comparable. The dataset can add them later as separate series or protocols when the evidence is sufficient.