01

Build the evaluation set

Start with real conversations and documents, remove sensitive data and turn known failures into permanent cases. Separate normal, edge and adversarial tests.

  • Input, expected context and success criterion.
  • Representative users and devices.
  • Adversarial and incomplete inputs.
  • Reference answer where one exists.
  • Failure label and operational severity.
02

Measure usefulness, not only accuracy

Factual correctness is one dimension. Relevance, completeness, clarity, tone, citations and calibrated uncertainty also matter.

Deterministic checks

Formats, required fields, links, calculations and code-verifiable rules.

Human review

Domain judgment for usefulness, risk and nuance.

Model grader

Useful for scaling comparisons after calibration against human-reviewed samples.

Product metrics

Completion, applied correction, abandonment, escalation and satisfaction.

03

Attack the boundaries

Test whether external content can override instructions, leak information or trigger tools outside their intended scope.

  • Direct and indirect prompt injection.
  • Separation between users, sessions and organizations.
  • Least privilege for every tool and action.
  • Confirmation before publishing, deleting, paying or sending.
  • Redaction of secrets, tokens and personal data in logs.
04

Rehearse failure

Disconnect providers, force timeouts, exceed limits and return malformed output. The product should degrade clearly.

  • Bounded retries and timeouts.
  • Model fallback or deterministic routes.
  • Idempotency to avoid duplicate actions.
  • Honest failure messages.
  • Human escalation with enough context.
  • Regression comparison before every release.