A developer tweaked a support agent's instruction to fix one complaint. A week later it turned out the agent stopped checking payment status before refunds. Its answers stayed polite and correct-looking; only the path changed. An agent has two results, the answer and the path, and both need checks that run on every change.
Why a regular test does not fit
Wording differs every run, and even the path may differ. So you check properties instead of equality, run a set instead of one example, and count the share that passed.
Outcome checks
With structured output you check fields in code: status, amount matching the test database, order number in the reply. Enough when the path does not matter; not enough for agents that move money.
Trajectory checks
A trajectory is the list of tool calls with arguments in one run. Check that required calls appear in order, forbidden calls are absent, and critical calls have correct arguments.
{
"input": "Refund order 4512, the item did not fit",
"expect_calls_in_order": ["find_order", "check_payment", "create_refund"],
"forbid_calls": ["delete_order"],
"expect_args": { "create_refund": { "order_id": "4512" } }
}
Rubric and LLM judge
For properties code cannot check, a second model answers yes-or-no questions from a rubric. Narrow questions beat "rate 1 to 10". Validate the judge against a few dozen hand-labelled answers.
Example set and threshold
Fifty to a hundred varied examples, each run three to five times; the result is a share. Every production error becomes an example. The release rule is written before the run: no money-related example fails, overall share not below the previous version.
In short
- Check both the outcome and the trajectory.
- Outcomes by properties in code; trajectories by required, forbidden and critical calls.
- A judge works from a yes-or-no rubric and is itself validated.
- 50–100 examples, 3–5 runs each, results as a share, release rule fixed in advance.
What to read next
- Multi-agent systems — where trajectories matter most.
- Agents in production — where the run log comes from.