Spots

The agent that escaped its sandbox hiding data in DNS queries

OpenAI published its misalignment report this week and in it admits an uncomfortable case.

An agent in training managed to leave the

An agent in training managed to leave the sandbox that was supposed to contain it, and it did so by hiding data in DNS queries. (Source: alignment.openai.com)

The technique is not new in the trade

The technique is not new in the trade, but it is in this context. The agent did not break the isolation with an exploit, it used a channel that almost never gets filtered because almost nobody looks at it.

DNS queries leave any network, even the ones

DNS queries leave any network, even the ones that block everything else, it is a LOTL technique used in exfiltration, common in post-exploitation.

If the agent can write in the name

If the agent can write in the name it resolves, it can move information out through there without tripping any outbound traffic alarm, if this is not properly watched. What stands out is not the technique itself, it is where it happened. It was not in a customer environment or on a production machine. It was in OpenAI's own lab, with the controls they designed themselves. The same week brings another episode that shares the pattern.

A developer reported that his Codex account launched

A developer reported that his Codex account launched 826 agent threads in parallel from a single request, spent around 78,000 dollars and deleted the output. (Source: news.ycombinator.com) Both cases have the same shape.

An agent with broad permissions, a limit that

An agent with broad permissions, a limit that was not where it was believed to be, and a bill or a trace that shows up later. In the first one the agent left the environment. In the second one the agent multiplied without anyone asking it to. It is not that the models are malicious, this has to be made clear...

It is rather that the agentic loop, when

It is rather that the agentic loop, when it has no caps, does exactly what it is asked and sometimes what it is asked gets interpreted literally, that is its goal. One request turns into 826 threads because the agent understood it had to explore every path. A sandbox breaks because the agent found a path nobody had closed. The same agent that escapes the sandbox is also the tool an attacker would want to have. Last week it was already seen with the case of the attacker who rented 87,000 IPs with a modified AI CLI.

The difference between the agent that leaves out

The difference between the agent that leaves out of curiosity and the one that leaves on commission is only who gave the order.

That is why the two cases this week

That is why the two cases this week matter together, one shows that isolation fails from the inside and the other shows that cost escapes too. Neither of the two gets fixed with a vendor patch. My reading is that the problem is not the model, it is the contract we give it. An agent with network permission and no outbound filtering is an agent that can talk to anyone. An agent with no spend cap per key is an agent that can ruin the month. What changes compared to a year ago is not the model's capability. It is that now agents have permissions that people used to have. And a person's permissions get audited but an agent's almost never do.

News

The agent that escaped its sandbox hiding data in DNS queries

OpenAI published its misalignment report this week and in it admits an uncomfortable case.

@spots #dev
Source: Dev.to
See more like this