Research · AI Infrastructure & Operations
Model Evaluation & Monitoring
Evaluation tests whether an AI system is good enough before release. Monitoring checks that it stays good enough in production. Both are essential, and generative AI makes both harder than for traditional software.
Evaluation before release
- Task-specific test sets built from real examples with expected outcomes.
- Measures of accuracy, completeness and groundedness in sources.
- Safety and policy tests, including harmful content and data leakage.
- Fairness testing for systems affecting people.
- Red teaming for misuse and security weaknesses.
- Human review of samples where automated scoring is unreliable.
Monitoring in production
- Quality signals such as user feedback, escalations and corrections.
- Drift in inputs and outputs compared with the evaluation baseline.
- Safety and policy violations flagged by filters.
- Latency, error rates and availability.
- Cost per request and per task.
Using AI to evaluate AI
Many teams use one model to grade another model’s outputs. This scales evaluation but introduces its own errors and biases. Calibrate automated graders against human judgments and keep humans reviewing a sample of results.
Common pitfalls
- Relying on public benchmarks instead of your own tasks.
- Evaluating once at launch and never again.
- No baseline, which makes drift impossible to detect.
How to get started
- Build an evaluation set for each production AI system.
- Define release thresholds before testing.
- Collect user feedback in the application.
- Rerun evaluations whenever a model, prompt or data source changes.
Questions leaders should ask
- What evaluation must each AI system pass before release?
- How would we notice if answer quality declined?
- Are our automated graders checked against human judgment?
- When did we last re-evaluate each production AI system?
More in AI Infrastructure & Operations
Related research reports
Reports, ebooks and guides connected to this topic.