Before saying one model wins, check what actually supports the ranking.
Two models can differ by a fraction of a point while the result still depends on benchmark composition, a small set of items, a nearly saturated ceiling or test exposure during training. This explorer keeps those failure modes separate instead of compressing them into an opaque reliability score.
—
Resolution and saturation
Aggregate scores do not reveal which exact items each model got right. We therefore do not infer paired significance or manufacture a p-value.
Dataset integrity
These are worst-case envelopes: they ask whether that number of items could explain the lead if every one favored the leader. They are not evidence of contamination or broken scoring.
Composition sensitivity
The calculation evaluates the extreme ± combinations for the selected percentage across all four task families. A winner flip shows dependence on which tasks we decide to emphasize.
Triggered fragilities
We do not add these signals into a single “reliability index.” Saturation, contamination, statistical resolution and composition are different problems that require different evidence.
What this explorer measures — and what it does not
A 95% Wilson interval is computed from the aggregate score and usable items. The approximation assumes binary scoring and that weights represent item shares. If task families use different metrics, or both models answer the same items, per-item data and an analysis matched to the evaluation design are needed.
The score gap is converted into an approximate number of items and compared with a hypothetical exposure. This is a sensitivity bound, not contamination detection.
Reweighting task families shows whether ranking depends on task selection. There is no universally correct weighting outside the deployment objective.
Accuracy alone does not cover robustness, calibration, efficiency, safety or deployment fit. This tool evaluates benchmark ranking claims, not total model quality.
A leaderboard needs statistical and methodological context.
The Benchmark Lottery shows that changing tasks can alter rankings. HELM argues for broad coverage, multiple metrics and explicit recognition of what is missing. LiveBench addresses contamination directly through frequently updated questions and objective scoring.