Build the evaluation set
Start with real conversations and documents, remove sensitive data and turn known failures into permanent cases. Separate normal, edge and adversarial tests.
- ✓Input, expected context and success criterion.
- ✓Representative users and devices.
- ✓Adversarial and incomplete inputs.
- ✓Reference answer where one exists.
- ✓Failure label and operational severity.
Measure usefulness, not only accuracy
Factual correctness is one dimension. Relevance, completeness, clarity, tone, citations and calibrated uncertainty also matter.
Deterministic checks
Formats, required fields, links, calculations and code-verifiable rules.
Human review
Domain judgment for usefulness, risk and nuance.
Model grader
Useful for scaling comparisons after calibration against human-reviewed samples.
Product metrics
Completion, applied correction, abandonment, escalation and satisfaction.
Attack the boundaries
Test whether external content can override instructions, leak information or trigger tools outside their intended scope.
- ✓Direct and indirect prompt injection.
- ✓Separation between users, sessions and organizations.
- ✓Least privilege for every tool and action.
- ✓Confirmation before publishing, deleting, paying or sending.
- ✓Redaction of secrets, tokens and personal data in logs.
Rehearse failure
Disconnect providers, force timeouts, exceed limits and return malformed output. The product should degrade clearly.
- ✓Bounded retries and timeouts.
- ✓Model fallback or deterministic routes.
- ✓Idempotency to avoid duplicate actions.
- ✓Honest failure messages.
- ✓Human escalation with enough context.
- ✓Regression comparison before every release.