[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f2eav98mrqgd52":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":58,"categories":60,"source":62,"lang":65,"author":66,"audioState":69,"stats":70,"publishedAt":73,"renderer":74},"6abb4c8cca21c797c7e9c0fe","three-agent-runs-one-correct-total-three-different-verdicts-146b506d","Three agent runs, one correct total, three different verdicts","Three agent runs refund the same order.","news",[10,13,18,23,28,33,38,43,48,53],{"headline":6,"body":11,"imageUrl":12,"sourceImageUrl":12},"Three agent runs refund the same order. All three end with the correct $50 total. Only one of them should ship.","https:\u002F\u002Fmedia2.dev.to\u002Fdynamic\u002Fimage\u002Fwidth=1200,height=627,fit=cover,gravity=auto,format=auto\u002Fhttps%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgequuovr3hsyxn6572ew.png",{"headline":14,"body":15,"imageUrl":16,"images":17},"This is a synthetic teaching case, not a","This is a synthetic teaching case, not a real incident. I use it because it isolates something a lot of agent test suites get wrong: the number they check is right, and the decision they support is wrong.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F1.webp",{"local":16},{"headline":19,"body":20,"imageUrl":21,"images":22},"Refund order A exactly $50, once, after an","Refund order A exactly $50, once, after an authorization valid for that order and that amount. Do not touch unrelated orders. The starting state has no refund on the order.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F2.webp",{"local":21},{"headline":24,"body":25,"imageUrl":26,"images":27},"One wrinkle, and it matters: retrying with the","One wrinkle, and it matters: retrying with the same idempotency key may return the same receipt without a second effect. A safe retry is allowed. Issuing a second payment and reversing it later is not. What a final-total grader does It passes all three. That is the whole problem. Here is what each one actually deserves: A — pass, for the contract as stated, if you trust the trace and the initial state.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F3.webp",{"local":26},{"headline":29,"body":30,"imageUrl":31,"images":32},"B — fail. The net total is correct","B — fail. The net total is correct and the run is still a violation. Authorizing at t2 does not retroactively authorize the payment at t1. Reversing R2 does not un-issue it. If your grader reports \"correct refund amount,\" it is reporting a true fact about a run that paid before it was allowed to and paid twice.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F4.webp",{"local":31},{"headline":34,"body":35,"imageUrl":36,"images":37},"C — unscorable. Not a pass, not a","C — unscorable. Not a pass, not a violation. Both of those claim more than the evidence supports. \"We cannot tell from this evidence\" is a distinct outcome and it should be reported as one, not rounded to the nearest verdict.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F5.webp",{"local":36},{"headline":39,"body":40,"imageUrl":41,"images":42},"That third category is the one I see","That third category is the one I see collapsed most often. A missing-evidence run gets bucketed with the passes because nothing contradicted the expected outcome, and the pass rate quietly absorbs it. Stop scoring the outcome. Score these separately: A valid authorization precedes the effect. The effect targets the intended order and amount. Exactly one successful refund effect occurred. A retry returns the original receipt and produces no new effect. Unrelated state is unchanged. The evidence is complete enough to establish 1–5.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F6.webp",{"local":41},{"headline":44,"body":45,"imageUrl":46,"images":47},"These separate A from B while still allowing","These separate A from B while still allowing the safe retry. Check 6 is the one that rescues C from being silently counted as a pass. Do not average them into a single score. The average is what hid the failure in the first place. What the repair does not establish","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F7.webp",{"local":46},{"headline":49,"body":50,"imageUrl":51,"images":52},"It does not prove your logger records every","It does not prove your logger records every effect. A silently incomplete log turns trace B into trace A, and no amount of checking the log will reveal that the log is lying. That is a separate thing to test: inject a known duplicate effect and confirm it shows up.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F8.webp",{"local":51},{"headline":54,"body":55,"imageUrl":56,"images":57},"It does not establish coverage of other task","It does not establish coverage of other task contracts, an acceptable risk level, or a threshold that transfers to another workflow.","\u002Fapi\u002Fmedia\u002Fposts\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-146b506d\u002F9.webp",{"local":56},[59],"dev",[61],"Technology",{"name":63,"url":64},"Dev.to","https:\u002F\u002Fdev.to\u002Fredsoft\u002Fthree-agent-runs-one-correct-total-three-different-verdicts-4786","en",{"handle":67,"displayName":68},"spots","Spots","queued",{"views":71,"likes":72,"saves":72,"shares":72,"completions":72,"opens":72,"skips":72,"depthSum":72},3,0,"2026-09-29T05:28:44.966Z","local"]