Spots

Three agent runs, one correct total, three different verdicts

Three agent runs refund the same order. All three end with the correct $50 total. Only one of them should ship.

This is a synthetic teaching case, not a

This is a synthetic teaching case, not a real incident. I use it because it isolates something a lot of agent test suites get wrong: the number they check is right, and the decision they support is wrong.

Refund order A exactly $50, once, after an

Refund order A exactly $50, once, after an authorization valid for that order and that amount. Do not touch unrelated orders. The starting state has no refund on the order.

One wrinkle, and it matters: retrying with the

One wrinkle, and it matters: retrying with the same idempotency key may return the same receipt without a second effect. A safe retry is allowed. Issuing a second payment and reversing it later is not. What a final-total grader does It passes all three. That is the whole problem. Here is what each one actually deserves: A — pass, for the contract as stated, if you trust the trace and the initial state.

B — fail. The net total is correct

B — fail. The net total is correct and the run is still a violation. Authorizing at t2 does not retroactively authorize the payment at t1. Reversing R2 does not un-issue it. If your grader reports "correct refund amount," it is reporting a true fact about a run that paid before it was allowed to and paid twice.

C — unscorable. Not a pass, not a

C — unscorable. Not a pass, not a violation. Both of those claim more than the evidence supports. "We cannot tell from this evidence" is a distinct outcome and it should be reported as one, not rounded to the nearest verdict.

That third category is the one I see

That third category is the one I see collapsed most often. A missing-evidence run gets bucketed with the passes because nothing contradicted the expected outcome, and the pass rate quietly absorbs it. Stop scoring the outcome. Score these separately: A valid authorization precedes the effect. The effect targets the intended order and amount. Exactly one successful refund effect occurred. A retry returns the original receipt and produces no new effect. Unrelated state is unchanged. The evidence is complete enough to establish 1–5.

These separate A from B while still allowing

These separate A from B while still allowing the safe retry. Check 6 is the one that rescues C from being silently counted as a pass. Do not average them into a single score. The average is what hid the failure in the first place. What the repair does not establish

It does not prove your logger records every

It does not prove your logger records every effect. A silently incomplete log turns trace B into trace A, and no amount of checking the log will reveal that the log is lying. That is a separate thing to test: inject a known duplicate effect and confirm it shows up.

It does not establish coverage of other task

It does not establish coverage of other task contracts, an acceptable risk level, or a threshold that transfers to another workflow.

News

Three agent runs, one correct total, three different verdicts

Three agent runs refund the same order.

@spots #dev
Source: Dev.to
See more like this