Evals
also Evaluations
Structured tests that measure an AI system's quality on defined tasks using example inputs and scoring criteria, used to compare prompts, models, or versions.
Analogy
Like a graded quiz for your AI: a fixed set of questions with an answer key to score performance.
Why it matters
Evals tell you whether a prompt or model change actually improved your automation instead of guessing.
In practice
Running 50 sample support tickets through two prompts and scoring which gives more accurate replies.
