Eval sets: curation, hard negatives, leakage, and versioning
A trustworthy eval set preserves provenance, separates banks by role, and freezes a release; hard pairs and distinct leakage channels test different failures without mutating the comparison.
Links containing ?t= open the video at a specific second.
Video summary
The ideas to retain
01
The object is an eval release, not "the dataset"
Represent one evaluation release as:
02
The smallest unit must be auditable
A case should let us resolve at least:
03
Start with a coverage hypothesis
"Representative" does not mean "random" by default.