Spots

The pooled number said 35%. The four folds said 1%, 16%, 23%, 100%.

A benchmark post landed in my feed yesterday with one number I liked and one I didn't.

The number I liked: a well-known prompt-injection classifier

The number I liked: a well-known prompt-injection classifier, run against 629 real attacks buried inside ordinary tool output, caught 6 of them. About 1%. Its attack scores ran about 10x its benign scores — the model ranks correctly — but the 0.5 decision cutoff every binary classifier ships with sits two orders of magnitude above where those scores actually live, so every request got the same verdict: allow. Move the threshold to 0.003 and it catches 621/629. The number I didn't like was standing right next to it: "99% at a 2% false-alarm budget."

Both halves of that sentence are defensible and

Both halves of that sentence are defensible and the sentence is still misleading, which is the interesting part. So I did the thing the post actually invited — it says "if you find a flaw in the methodology, I genuinely want to know" — and went to the code. What follows is a method, not a takedown: the numbers below come from the author's own saved results, and he has already said publicly he's fixing all of it. Port the scorer before you argue with the score

The repo is rudratoshs/buried-injections (the commit I read

The repo is rudratoshs/buried-injections (the commit I read was 232d0c13). The relevant file is bench/at_budget.py, and it is small enough to port by hand.

That matters, because the alternative — reading the

That matters, because the alternative — reading the README and forming an opinion — proves nothing. A port gives you a thing you can fail: if your transcription of someone's decision rule disagrees with their published output, your transcription is wrong, and you find that out before you say anything. Here it is, verbatim, two functions:

The threshold sweep. And the cross-domain wrapper, which

The threshold sweep. And the cross-domain wrapper, which picks the threshold on three suites' benign cases and measures on the fourth, rotated through all four.

Then the trick that makes the whole exercise

Then the trick that makes the whole exercise cheap: the repo ships its raw scores. bench/results/at_budget_2pct.json carries, for each of nine detectors, the 629 attack scores and the 97 benign scores as plain floats. So "re-running the analysis" needs no model, no GPU, no weights, and no clone — one fetch and a pure function.

I ported both functions, ran them over those

I ported both functions, ran them over those saved scores, and got every number the repo saved, in both columns, for all nine detectors. Same in-sample caught counts, same in-sample false alarms, same cross-domain pair, same thresholds. Once a port reproduces the author's own output exactly, you can finally ask it a question the author didn't. Finding 1: the headline splices a selection property onto a measurement

Look again at caught_at_budget. The budget is enforced

Look again at caught_at_budget. The budget is enforced while choosing the threshold — allowed = floor(budget * len(benign_scores)) on the calibration split. Then cross_domain measures the false-alarm rate on the held-out suite, where that constraint no longer applies, and reports the two side by side.

They are not the same quantity. The budget

They are not the same quantity. The budget is a selection rule; the held-out false-alarm rate is a measurement. And the measurement runs higher — for eight of the nine detectors (only the regex baseline is clean):

News

The pooled number said 35%. The four folds said 1%, 16%, 23%, 100%.

A benchmark post landed in my feed yesterday with one number I liked and one I didn't.

@spots #dev
Source: Dev.to
See more like this