A developer tweaked a support agent's instruction to fix one complaint. A week later it turned out the agent stopped checking payment status before refunds. Its answers stayed polite and correct-looking; only the path changed. An agent has two results, the answer and the path, and both need checks that run on every change.

Why a regular test does not fit

Wording differs every run, and even the path may differ. So you check properties instead of equality, run a set instead of one example, and count the share that passed.

Outcome checks

With structured output you check fields in code: status, amount matching the test database, order number in the reply. Enough when the path does not matter; not enough for agents that move money.

Trajectory checks

A trajectory is the list of tool calls with arguments in one run. Check that required calls appear in order, forbidden calls are absent, and critical calls have correct arguments.

{
  "input": "Refund order 4512, the item did not fit",
  "expect_calls_in_order": ["find_order", "check_payment", "create_refund"],
  "forbid_calls": ["delete_order"],
  "expect_args": { "create_refund": { "order_id": "4512" } }
}

Rubric and LLM judge

For properties code cannot check, a second model answers yes-or-no questions from a rubric. Narrow questions beat "rate 1 to 10". Validate the judge against a few dozen hand-labelled answers.

Example set and threshold

Fifty to a hundred varied examples, each run three to five times; the result is a share. Every production error becomes an example. The release rule is written before the run: no money-related example fails, overall share not below the previous version.

In short

  • Check both the outcome and the trajectory.
  • Outcomes by properties in code; trajectories by required, forbidden and critical calls.
  • A judge works from a yes-or-no rubric and is itself validated.
  • 50–100 examples, 3–5 runs each, results as a share, release rule fixed in advance.