Tools · Evaluation · 13

Before saying one model wins, check what actually supports the ranking.

Two models can differ by a fraction of a point while the result still depends on benchmark composition, a small set of items, a nearly saturated ceiling or test exposure during training. This explorer keeps those failure modes separate instead of compressing them into an opaque reliability score.

Resolutionscore intervals
Integrityinvalid items and exposure
Compositionweight sensitivity
Boundarybenchmark ≠ total quality

Evaluation set

Binary item scoring is assumed for the descriptive interval.
Broken, ambiguous or incorrectly referenced questions.
A sensitivity scenario; it does not claim contamination exists.
Tests how strongly the winner depends on benchmark composition.

Tasks and scores

TaskWeightModel AModel BGap
Reasoning
Coding
Knowledge
Instruction following

Weights are normalized automatically. Interpret them as the share of benchmark items represented by each task family.

Model A
Model B
Observed difference items of approximate lead
Usable items

Resolution and saturation

95% interval overlapdescriptive diagnostic, not a paired test
Headroom to 100%less headroom means higher saturation risk

Aggregate scores do not reveal which exact items each model got right. We therefore do not infer paired significance or manufacture a p-value.

Dataset integrity

Estimated invalid itemscover the gap?
Potentially exposed itemscover the gap?

These are worst-case envelopes: they ask whether that number of items could explain the lead if every one favored the leader. They are not evidence of contamination or broken scoring.

Composition sensitivity

Gap range after reweightingminimum → maximum
Rankingall weights are renormalized

The calculation evaluates the extreme ± combinations for the selected percentage across all four task families. A winner flip shows dependence on which tasks we decide to emphasize.

Triggered fragilities

We do not add these signals into a single “reliability index.” Saturation, contamination, statistical resolution and composition are different problems that require different evidence.

Method

What this explorer measures — and what it does not

Descriptive interval

A 95% Wilson interval is computed from the aggregate score and usable items. The approximation assumes binary scoring and that weights represent item shares. If task families use different metrics, or both models answer the same items, per-item data and an analysis matched to the evaluation design are needed.

Contamination sensitivity

The score gap is converted into an approximate number of items and compared with a hypothetical exposure. This is a sensitivity bound, not contamination detection.

Benchmark lottery

Reweighting task families shows whether ranking depends on task selection. There is no universally correct weighting outside the deployment objective.

Holistic evaluation

Accuracy alone does not cover robustness, calibration, efficiency, safety or deployment fit. This tool evaluates benchmark ranking claims, not total model quality.

Sources

A leaderboard needs statistical and methodological context.

The Benchmark Lottery shows that changing tasks can alter rankings. HELM argues for broad coverage, multiple metrics and explicit recognition of what is missing. LiveBench addresses contamination directly through frequently updated questions and objective scoring.