The Eval Set Is the Product
Models get replaced. The labeled test set that judges them outlives every one of them, which is why it deserves more care than the training run.
Every model I ship gets replaced. That’s the schedule. A better backbone comes out, a client’s data distribution shifts, a vendor changes pricing and I swap providers, or I just find an architecture that gets the same recall for a third of the compute. The model in production today is a rental.
I’ve long believed most AI problems can be fixed with simple data quality rules, and the eval set is where that belief matters most. It’s the one thing that doesn’t rotate out: the fixed, labeled, held-out data that answers whether the new model is better than the old one on the cases that matter. It answers that the same way regardless of which framework or vendor happens to be running this quarter. Treat it as an afterthought to the training run and you’ve built your whole measurement system on sand, then poured an expensive model on top of it.
Label quality matters more on the eval set than the training set
Training data absorbs label noise fine. A percent or two of mislabeled examples averages out over thousands of gradient steps, and a model can often outperform the noisy labels it learned from. Eval labels don’t get that allowance. Every one is a claim about ground truth, and every wrong claim either hides a real regression or manufactures a fake one.
If 2% of your training labels are wrong, you probably never notice. If 2% of your eval labels are wrong, that’s a noise floor under every comparison you run against that set. Ship a model that’s really 1% better and the eval set can’t see it. Worse, once that’s suspected, a team under deadline pressure starts waving away real regressions as “probably a label error,” and now nobody trusts the number in either direction. A CV eval set with sloppy bounding boxes or ambiguous class boundaries adds noise that’s indistinguishable from a real regression, and that corrodes trust in the whole measurement.
The fix is discipline: adjudicate eval labels more rigorously than training labels, get a second reviewer on anything ambiguous, and write the labeling guideline down so “ambiguous” means the same thing in month six that it meant in month one.
Stratify by the failure modes the business cares about
An eval set that’s a random sample of production traffic tells you average-case accuracy and little else. That’s rarely the number worth optimizing, because the failures that hurt a business cluster in the tail, and a random sample under-represents the tail by construction. Build the set instead around the specific ways a model can fail that cost something:
- Rare classes or edge cases that are individually infrequent but collectively define trustworthiness, where wrong answers are expensive even if uncommon.
- Known hard cases (poor lighting, occlusion, unusual angles, degraded input), pulled into their own slice rather than diluted into an aggregate score.
- Near-boundary examples, deliberately chosen at the decision threshold where small model changes flip the answer.
Report accuracy per slice as well as in aggregate. A model can improve its overall number while getting worse on the rare-but-expensive slice, and an aggregate metric hides exactly that trade the whole way to production. This is the same argument I make about scope in AI integration work generally: the value is in knowing which failures are tolerable, and an unstratified eval set throws that knowledge away before it can inform anything.
Tip
When you add a new slice to the eval set, backfill scores for every model you still have artifacts for. A slice with one data point tells you nothing about whether it discriminates between models or just reflects one model’s quirks.
Version it, never train on it, and don’t let it belong to the vendor
The eval set needs the same version discipline as the code consuming it (a hash or tag, checked in with a note on when and why it changed) so “we improved 3 points” traces back to a specific, frozen set of examples rather than a moving target that got easier. Compare model A on eval-v3 against model B on eval-v4 and you’re comparing eval sets, not models, and nobody says so out loud when the numbers hit a slide.
The harder discipline is the wall between eval and training. Everyone agrees with “never train on your test set” in principle, then it happens anyway through the side door: a labeler recognizes an eval image and adds a similar one to the training pool; a fine-tune draws on a broader internal dataset assembled before the eval set was carved out, overlap unchecked; a near-duplicate frame from the same video ends up in training because nobody deduplicated across the split. Once that happens, every number the eval set produces afterward is contaminated while still looking exactly as trustworthy as before. That’s the worst failure mode available, because it never announces itself. Deduplication has to be an explicit, checked step every time new training data comes in.
And regardless of who trains the model (me, another consultancy, an in-house team eighteen months from now), the eval set should belong to the client outright: their examples, their labels, their stratification, sitting in their infrastructure, independent of whoever’s contract happens to be active. A vendor who trains the model and also controls the eval set is grading their own homework, and even an honest one will unconsciously shape the set toward the cases their approach handles well. That’s not a claim about anyone’s integrity. It’s what happens when the measuring stick and the thing being measured share an owner. Client ownership of the eval set lets them compare vendors honestly, catch silent regressions when a model gets swapped, and keep a permanent record of what “better” has meant over the life of the system.
Models, frameworks, and eventually even the engineer will change. A well-built, well-labeled, properly stratified eval set is the one artifact from a computer vision engagement that should still be answering the same question, the same way, five model generations later.
Content may be edited with the help of AI. All content has been reviewed by a human before being published.
Have a project in mind?
I design and ship custom software — from early concept to production. Tell me what you're building and we'll figure it out together.