Spots

Checkpoint Restore: Destination Policy Was Never Reapplied

A checkpoint restore can rebuild a process from saved state without passing that state back through the destination's normal policy translation. When it does, the security context on the Pod spec describes the workload that was approved, and the process on the node is the one that was saved.

That is a boundary problem before it is

That is a boundary problem before it is a vulnerability problem. Kubernetes has one place where policy becomes process state, and it is creation. Restore is a second route to a running process, and nothing on that route decides what wins when saved state and destination policy disagree. Advisories on the restore path in containerd and CRI-O show what crosses when nothing decides. They are evidence for the architecture, and the architecture is the subject. Two Ways to Create a Process

On the normal path, the process is built

On the normal path, the process is built from declarations. A Pod spec passes admission, the kubelet hands the runtime a container config, and the runtime turns that config into kernel state: a user and group, a capability set, the no_new_privs flag, a seccomp filter. That translation step is how modern infrastructure architecture turns declared intent into enforced state. Each attribute is policy made concrete, which is why seccomp filters and capability drops eliminate whole classes of exploit only when they are actually attached to the process. Kubernetes documents that allowPrivilegeEscalation directly controls whether no_new_privs is set. The field in the spec is the request. The flag on the running process is the enforcement.

Checkpoint restore does not build a process. It

Checkpoint restore does not build a process. It reconstitutes one. CRIU records a running process, including its credentials, capabilities, no_new_privs flag and seccomp state, and restore replays that record. On the path containerd exposed, the restore is triggered inside the ordinary container-create call, by a checkpoint archive or an annotated image. Kubernetes currently supports container restore only through those image annotations, so from admission's side this is a Pod create with an image reference. Whether it is a restore is decided later, inside the runtime, by what the image contains.

Restore is usually treated as a data operation

Restore is usually treated as a data operation whose failures show up in the layers above it, the pattern restore design failure documents: the restore completes and the workload is still unusable. This is a different gap. Every layer can pass, and the process can still come back carrying state the destination never approved. What the Advisories Show About Checkpoint Restore

containerd's September 1 advisory states the outcome directly

containerd's September 1 advisory states the outcome directly. A container restored from an untrusted checkpoint through the create call can run as root with full capabilities and no seccomp filter, despite restrictive policy requested by the orchestrator. It applies where checkpoint restore through CRI is enabled and an attacker can run a container from a crafted checkpoint image. The affected ranges run from 2.1.0 up to, but not including, 2.2.7, and from 2.3.0 up to, but not including, 2.3.4. Those two releases disable the path by default.

The same checkpoint and restore path had already

The same checkpoint and restore path had already produced three other trust failures in June, according to Google's GKE security bulletins. Each accepted something the checkpoint supplied.

CRI-O's CVE-2026-92574 records the same outcome for that

CRI-O's CVE-2026-92574 records the same outcome for that runtime: saved credentials, capabilities, no_new_privs and seccomp state can take effect where the destination's configuration should have.

⚠ Scope of exposure: Each of these needs

⚠ Scope of exposure: Each of these needs the ability to create Pods, and the September issue matters only where checkpoint restore is enabled. containerd rates it Critical. Google rates it Medium for GKE and notes that default GKE nodes ship without the criu binary. The point here is the boundary, not the count of exposed clusters.

⚠ Scope of exposure: Each of these needs

⚠ Scope of exposure: Each of these needs the ability to create Pods, and the September issue matters only where checkpoint restore is enabled. containerd rates it Critical. Google rates it Medium for GKE and notes that default GKE nodes ship without the criu binary. The point here is the boundary, not the count of exposed clusters. Precedence Was Never Decided

News

Checkpoint Restore: Destination Policy Was Never Reapplied

A checkpoint restore can rebuild a process from saved state without passing that state back through the destination's normal policy translation.

@spots #dev
Source: Dev.to
See more like this