Traditional software is tested with expected outputs. Generative AI produces different wording every time, which tempts teams to skip rigorous testing and rely on demos. That is the single biggest reason GenAI projects disappoint in production. Evaluation is how you replace opinions with evidence.
Step 1: Build an evaluation set
Collect 50–300 real examples of the inputs your application will receive, with the expected answer or the key facts a good answer must contain. Include easy cases, hard cases, edge cases and inputs the system should refuse. Involve the business experts who will judge quality in production.
Step 2: Choose the right metrics
| What to measure | Examples |
|---|---|
| Correctness | Does the answer contain the required facts? Is it free of errors? |
| Groundedness | Is every claim supported by the retrieved sources? |
| Retrieval quality (RAG) | Were the right documents retrieved? Recall and precision of context |
| Safety | Refuses harmful or out-of-scope requests; no sensitive-data leakage |
| Format & tone | Follows the required structure and style |
| Operational | Latency, cost per request, error rate |
Step 3: Combine automated and human evaluation
- Deterministic checks — format validation, required keywords, citation presence — are cheap and reliable.
- LLM-as-judge — using a model to grade answers against criteria — scales well, but must itself be validated against human judgements and kept consistent.
- Human review — essential for a calibration sample and for high-stakes use cases.
Step 4: Run evaluations continuously
Treat prompts, retrieval settings and model versions like code. Run the evaluation suite in CI/CD on every change and block releases that drop below agreed thresholds. Re-run when the underlying model changes — even minor model updates can shift behaviour.
Step 5: Monitor in production
- Capture user feedback (thumbs up/down, edits made to outputs).
- Sample production conversations for review and add failures to the evaluation set.
- Watch latency, cost and refusal rates for drift.
Common mistakes
- Testing only on a handful of demo questions.
- Trusting an LLM judge without checking it against human ratings.
- Measuring model quality but not retrieval quality in RAG systems.
- Never re-testing after model or data changes.
Our Generative AI team sets up evaluation frameworks and LLMOps pipelines as part of every engagement.
Frequently asked questions
How many test cases do I need to evaluate a generative AI app?
Many teams start with 50 to 300 representative examples covering easy, hard, edge and refusal cases, growing the set over time with real production failures.
What is LLM-as-judge?
It is the practice of using a language model to grade another model's outputs against defined criteria. It scales well but should be validated against human judgements.
How do you evaluate a RAG system?
Measure retrieval quality (were the right sources retrieved?), groundedness (are claims supported by those sources?) and answer correctness, plus latency and cost.