
Three agent runs, one correct total, three different verdicts
Three agent runs refund the same order. All three end with the correct $50 total. Only one of them should ship.
Spots 

Three agent runs refund the same order. All three end with the correct $50 total. Only one of them should ship.

This is a synthetic teaching case, not a real incident. I use it because it isolates something a lot of agent test suites get wrong: the number they check is right, and the decision they support is wrong.
Refund order A exactly $50, once, after an authorization valid for that order and that amount. Do not touch unrelated orders. The starting state has no refund on the order.
One wrinkle, and it matters: retrying with the same idempotency key may return the same receipt without a second effect. A safe retry is allowed. Issuing a second payment and reversing it later is not. What a final-total grader does It passes all three. That is the whole problem. Here is what each one actually deserves: A — pass, for the contract as stated, if you trust the trace and the initial state.
B — fail. The net total is correct and the run is still a violation. Authorizing at t2 does not retroactively authorize the payment at t1. Reversing R2 does not un-issue it. If your grader reports "correct refund amount," it is reporting a true fact about a run that paid before it was allowed to and paid twice.
C — unscorable. Not a pass, not a violation. Both of those claim more than the evidence supports. "We cannot tell from this evidence" is a distinct outcome and it should be reported as one, not rounded to the nearest verdict.
That third category is the one I see collapsed most often. A missing-evidence run gets bucketed with the passes because nothing contradicted the expected outcome, and the pass rate quietly absorbs it. Stop scoring the outcome. Score these separately: A valid authorization precedes the effect. The effect targets the intended order and amount. Exactly one successful refund effect occurred. A retry returns the original receipt and produces no new effect. Unrelated state is unchanged. The evidence is complete enough to establish 1–5.
These separate A from B while still allowing the safe retry. Check 6 is the one that rescues C from being silently counted as a pass. Do not average them into a single score. The average is what hid the failure in the first place. What the repair does not establish
It does not prove your logger records every effect. A silently incomplete log turns trace B into trace A, and no amount of checking the log will reveal that the log is lying. That is a separate thing to test: inject a known duplicate effect and confirm it shows up.
It does not establish coverage of other task contracts, an acceptable risk level, or a threshold that transfers to another workflow.
Three agent runs refund the same order.
