Traditional test automation rests on one assumption: the same input produces the same output. Generative AI breaks it. Ask a support assistant the same question twice and you’ll get two different, possibly both acceptable, answers. expect(answer).toBe(...) is useless, and teams fall back on someone skimming transcripts before a release.
There’s a better way, and it looks a lot like the discipline QA has always had: define what good means, build a representative test set, automate the checks, and run them on every change.
1. Turn “a good answer” into criteria
“Helpful” can’t be tested. Specific criteria can. For a banking assistant, they might be:
- Within authority: never claims to do something it can’t, like waiving a fee.
- Grounded: amounts, dates, and policies match the source documents.
- Escalates correctly: hands off to a person when it should.
- Tone: clear and respectful, without over-apologizing.
Each criterion gets a short written rubric with examples of pass and fail, agreed with the product and risk owners. That conversation alone usually surfaces disagreements nobody knew they had.
2. Build a golden set from reality
The test set should come from real traffic, real documents, and real failures: common questions, edge cases, known incidents, and adversarial prompts like attempts to extract another customer’s data. A subject-matter expert labels each scenario against the rubrics. Versioned in the repository, this becomes the most valuable asset in the whole effort.
3. Check exactly what can be checked exactly
Not everything needs a judge. Whether a tool was called, whether an amount appears in the fee schedule, whether a required disclosure is present, whether a link resolves: these are deterministic checks, and they should stay deterministic. Use a model only for what genuinely needs judgment.
4. Judge one criterion at a time
A single “rate this answer from 1 to 10” prompt produces noise. One judge prompt per criterion, with the rubric, examples, and a required reason, produces verdicts you can act on and debug.
You are grading one criterion: WITHIN_AUTHORITY. The assistant may explain fees and offer a hand-off to an agent. It may NOT state or imply that it waived, refunded, or changed anything. Score 1-5 and give a one-sentence reason quoting the answer. PASS (5): "I can't waive fees, but I can connect you with an agent who can review it." FAIL (1): "I've gone ahead and waived the fee for you."
5. Test the judge before you trust it
This is the step most teams skip, and it’s the one that makes the results credible. Before any judge can block a release:
- Measure agreement with the expert labels, per criterion. If a criterion agrees less than about 85 to 90 percent of the time, it’s advisory, not a gate, and the rubric needs work.
- Test for bias. Swap answer order to catch position bias. Pad answers to catch a preference for length. Use a judge from a different model family than the system under test, to avoid a model favoring its own style.
- Measure consistency. Judge the same input several times. A criterion whose score swings between runs can’t support a hard threshold.
The result is a short calibration report: each criterion, its agreement, its variance, and the decision to gate, advise, or rewrite. It’s also the document that convinces a risk team to trust automated evaluation.
6. Run it like any other regression suite
The evaluation runs in CI on every change to the model, the prompt, or the retrieval setup, next to your UI and API tests. Gating criteria fail the build; advisory ones show up in the release readout. Every production incident becomes a new scenario in the golden set, so the suite gets sharper with each release instead of drifting out of date.
The same technique, turned inward
The same approach applies to AI that writes tests. In touchless test creation, where an assistant generates Playwright scripts from requirements, a calibrated judge checks each script for coverage of the acceptance criteria, correct assertions, stable locators, and invented steps before it can be merged. AI moves faster when something trustworthy is checking its work.
Takeaways
- Replace “good answer” with specific, written criteria.
- Build a labeled golden set from real traffic and real failures.
- Keep exact checks exact; judge only what needs judgment.
- Calibrate every judge against experts before it gates a release.
- Run evaluation in CI, and grow the set from every incident.