Glossary · Updated Sep 2, 2026

Evaluation (evals)

Systematic testing of an AI system's outputs against expected results or quality criteria to measure and track performance.

Definition

Evaluations range from automated checks on curated test sets to human ratings and model-graded rubrics. Teams run them when changing prompts, models, or retrieval settings to catch regressions, and vendors use them to publish benchmarks. Good evals reflect real user tasks rather than generic academic tests.

Why it matters when choosing a tool

Without evaluation you cannot know whether a tool or prompt change improved things; ask vendors how they measure quality on tasks like yours.

Where you will meet it

AI Coding & Development, AI Agent & Chatbot Builders, AI Customer Support

Related terms

Benchmark · Prompt engineering · Guardrails

Tools where this matters

Reviewed products in the related categories