Evaluation set
Representative normal, edge, adversarial, and failure cases used to test a workflow repeatedly.
Definition
Representative normal, edge, adversarial, and failure cases used to test a workflow repeatedly.
Practical test
Does the set represent normal, edge, adversarial, and known-failure cases in the real task distribution?
Example
A document extractor set includes clean files, blank fields, rotated scans, duplicate pages, handwriting, and conflicting totals.
What to record
For evaluation set to be operational rather than a reassuring label, record case provenance, expected result, covered requirement, reviewer, version, and observed result. If those facts cannot be observed or tested, do not treat the term itself as evidence that the workflow is safe.
