All videosOther topics0:36

Eval sets: curation, hard negatives, leakage, and versioning

A trustworthy eval set preserves provenance, separates banks by role, and freezes a release; hard pairs and distinct leakage channels test different failures without mutating the comparison.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

The object is an eval release, not "the dataset"

Represent one evaluation release as:

02

The smallest unit must be auditable

A case should let us resolve at least:

03

Start with a coverage hypothesis

"Representative" does not mean "random" by default.

Key moments

Jump directly to a section

  1. Provenance and grouping come before the split
  2. Hard positive and hard negative cross one boundary
  3. Freeze the release; new failures feed v+1