AJ Learning Hub logoAJ Learning Hub

Evals

also Evaluations

Structured tests that measure an AI system's quality on defined tasks using example inputs and scoring criteria, used to compare prompts, models, or versions.

Analogy

Like a graded quiz for your AI: a fixed set of questions with an answer key to score performance.

Why it matters

Evals tell you whether a prompt or model change actually improved your automation instead of guessing.

In practice

Running 50 sample support tickets through two prompts and scoring which gives more accurate replies.

Related terms:GuardrailsHallucinationSystem PromptAgent