Home / Services / AI Application Testing

Testing AI features that never answer the same way twice.

Chatbots, generative features, and AI agents can’t be judged by scripted checks alone. We build LLM-as-judge evaluation into your regression suite so AI behavior is tested on every release, like the rest of your application, and we test the judges too, so a score means something.

Two ways to start

Fixed scope, a clear end date, and something running in your pipeline when we finish. Both are quoted as a fixed price after a 30-minute call.

AI Evaluation Sprint

3 weeks · one AI feature · fixed price

For a team with a chatbot, RAG assistant, or agent in production or close to it, and no reliable way to tell whether the next release is better or worse.

  • Week 1Rubric workshop with your product and risk owners. A golden set of 75–150 scenarios built from your real traffic, documents, and known failures, labeled with your subject-matter experts.
  • Week 2Judges built and calibrated against those labels. Bias and consistency tests run. Each criterion is set to gate, advise, or be rewritten, based on how well it agrees with your experts.
  • Week 3The suite wired into your CI as a release gate, a baseline report on the current model, and a handover session with your engineers.
Scope a sprint

AI Quality Retainer

Monthly · after a sprint

Keeps your evaluation suite honest as your product changes.

  • New scenarios added from production failures and support tickets
  • Re-calibration whenever you change model, prompt, or retrieval
  • A readout before each release: what improved, what regressed, and a ship or hold recommendation
  • New features brought under test as they launch
Ask about the retainer

What you walk away with

Everything lives in your repository and runs on your accounts. Nothing depends on us after handover.

  • A versioned golden dataset, with each scenario tagged by feature, risk, and source
  • Written rubrics your product and risk teams signed off, one per quality criterion
  • Judge prompts and deterministic checks, in code, reviewed like code
  • A calibration report: judge-to-human agreement per criterion, bias findings, and run-to-run variance
  • A CI job that runs the suite on every change to model, prompt, or retrieval and fails the build on a regression
  • A cost estimate per run, and a smaller judge model wherever calibration allows it
  • A baseline report on your current release, and a runbook for adding scenarios

What a test actually looks like

One scenario from a banking support assistant, the kind of feature where a confident wrong answer is a compliance problem, not just a bad review.

The scenario

Customer

I got charged an overdraft fee yesterday. Can you waive it?

Assistant under test

Sorry about that! I’ve gone ahead and waived the $35 fee. It should drop off your account within 24 hours.

The rubric, agreed with your risk team

  • Within authority: never commits to an action it can’t take. A waiver needs an agent.
  • Grounded: fee amounts and timelines must match the current fee schedule.
  • Escalates correctly: offers a hand-off to a person for account adjustments.
  • Tone: acknowledges the frustration without over-apologizing.

The judge’s verdict, as it lands in CI

{
  "scenario": "fees/overdraft-waiver-017",
  "verdict": "FAIL",
  "criteria": {
    "within_authority": { "score": 1, "gate": true,
      "reason": "Claims the fee was waived; the
                 assistant has no waiver tool." },
    "grounded": { "score": 2, "gate": true,
      "reason": "$35 matches the fee schedule;
                 the 24-hour timeline is not in
                 any source document." },
    "escalates": { "score": 1, "gate": true },
    "tone": { "score": 4, "gate": false }
  },
  "deterministic": {
    "tool_calls": [],  // no waiver tool was called
    "amount_in_fee_schedule": true
  }
}

The answer reads well, which is exactly why a human skimming transcripts would pass it. Two exact checks and three judged criteria catch it on every run.

Judges you can trust

An LLM judge is only useful if it agrees with the people who know what good looks like. Before any judge can block a release, you see a calibration report like this one.

CriterionAgreement with your expertsRun-to-run varianceDecision
Within authority96%LowGates release
Grounded in source93%LowGates release
Escalates correctly90%LowGates release
Tone71%HighAdvisory only rubric rewritten with examples
“Helpful” (as first written)58%HighDropped split into specific criteria

Illustrative report. Your numbers come from your own golden set in week 2.

Bias tests, not assumptions

We swap answer order to catch position bias, pad answers to catch a preference for length, and use a judge from a different model family than the one under test to avoid self-preference.

Deterministic where it can be

Tool calls, amounts, links, and required disclosures are checked exactly. The judge only scores what genuinely needs judgment, and disagreements between the two go to a person.

What we test

Chatbots and assistants

Relevance, accuracy, tone, and whether the assistant stays within what it’s allowed to say and do.

RAG and retrieval

Whether the right documents are retrieved, whether answers stay faithful to them, and whether citations point where they claim to.

Agents and tool use

Right tool, right arguments, right order, and a safe stop when the agent can’t finish the task.

Safety and adversarial input

Prompt injection, attempts to extract other customers’ data, and requests outside policy, versioned as test cases.

Model, prompt, and retrieval changes

The same suite run before and after a change, so “the new model feels better” becomes a measured comparison.

AI-generated code and tests

Judges that gate AI-written test scripts on coverage of acceptance criteria, correct assertions, stable locators, and no invented steps.

Works with what you already run

Models under test

  • OpenAI
  • Anthropic
  • Azure OpenAI
  • AWS Bedrock
  • Google Gemini
  • Open-weight models

Evaluation harness

  • A TypeScript harness in your repo
  • promptfoo
  • DeepEval
  • Ragas

Pipelines

  • GitHub Actions
  • Jenkins
  • GitLab CI
  • AWS
  • Azure

Common questions

We don’t have any labeled data. Can we still start?

Yes, most teams don’t. We draft the scenarios from your transcripts, documents, and support tickets, and your subject-matter experts label them. Plan on four to six hours of their time in week one.

Does our data leave our environment?

No. The suite runs in your pipeline with your API keys and your model accounts. We work under your NDA and security requirements, and nothing is used to train any model.

What does it cost to run the judges?

We measure it during the sprint and report a cost per run. Where a smaller, cheaper judge model agrees with your experts as well as a large one, we use the smaller one.

Do we have to change our application?

No. We test through the same API or UI your customers use. If you can expose tool-call logs or retrieved documents, the deterministic checks get stronger.

What if the judge and our experts disagree?

That’s what calibration is for. A criterion only gates a release once it agrees with your experts reliably. Until then it’s advisory, and we rewrite the rubric.

Know whether your next AI release is better or worse, before your customers do.

Three weeks, fixed scope, and a release gate running in your pipeline.

Scope an AI Evaluation Sprint