Evaluation foundations
Eval
A structured test of whether an AI system meets a requirement.
Also known as: evaluation
Definition
An eval is any repeatable assessment of an AI system against an expected behaviour, quality threshold, risk control or business outcome.
Why it matters
Unlike a public benchmark, an eval can be designed around your exact use case, documents, policies, users and failure tolerance.
Related concepts
- Benchmark
A standardised test used to compare systems consistently.
- Golden set
A curated set of representative inputs and trusted expected outcomes.
- System under test
The complete configuration being evaluated—not only the model.
- Continuous evaluation
Rerunning tests as a system and its environment change.