Spots

PQC-Bench Part 1 — measuring X25519MLKEM768 costs on a $0.04/hour Graviton3

There's a lot of discussion going on about post-quantum TLS. Both AWS and Cloudflare claim the overhead is minimal — but both are vendor-published, on infrastructure most people don't have. This is the starting post of a different attempt: an independent, reproducible benchmark of X25519MLKEM768 (the NIST-standardized hybrid PQ key exchange, FIPS 203) versus classical X25519. I ran it on two c7g.large EC2 instances in the same AWS zone. The full run cost less than a dollar.

Complete setup — Terraform, Gatling simulation, driver script

Complete setup — Terraform, Gatling simulation, driver script, raw Gatling result files — is in the repo. This can be reproduced in around ~50 minutes for under $0.30 (based on my experience, accounting for the occasional disruption during a test). Note on the self-signed cert The server serves an ECDSA cert, not an ML-DSA-65 one.

The reason: Gatling uses Netty's BoringSSL under the

The reason: Gatling uses Netty's BoringSSL under the hood, and BoringSSL doesn't advertise ML-DSA-65 in its supported signature algorithms. If nginx only had an ML-DSA-65 cert to serve, BoringSSL wouldn't be able to complete the handshake at all.

This actually mirrors real-world PQ TLS in 2026

This actually mirrors real-world PQ TLS in 2026: no public CA issues ML-DSA-65 certs. Everyone shipping "PQ TLS" today serves a classical cert with a hybrid KEM. This benchmark measures the KEM overhead under exactly that pattern. First attempt: what didn't work My first serious run was at 1000 req/s. The numbers were absurd:

Not "PQ is faster than classical" — that's

Not "PQ is faster than classical" — that's not physically possible when the PQ arm does the same work as classical plus more crypto. The shape (min=1ms, mean=1300ms) is the fingerprint of client-side CPU saturation. 1000 fresh handshakes/sec on a 2-vCPU loadgen means each vCPU is doing 500 handshakes/sec of X25519 or ML-KEM math — beyond a c7g.large's capacity. Requests queue for CPU, means shoot into the seconds range, and any actual PQ vs classical delta gets swallowed by queuing noise.

Cross-check: PQ p99 across the 3 trials ranged

Cross-check: PQ p99 across the 3 trials ranged 5892 → 6091 → 7623 ms. That trial-to-trial variance is bigger than any real signal I could measure — a red flag on its own.

Lesson: benchmark your load generator before you trust

Lesson: benchmark your load generator before you trust its numbers. Client-side CPU cost is easy to forget when the discussion is all about server overhead. The real numbers, at 300 req/s Dropping to 300 req/s puts the loadgen well inside its capacity envelope. Now I'm measuring the handshake itself. Median of 3 trials, latency in milliseconds: Trial-by-trial p99 (to show the variance honestly):

Trial 3's PQ p99 of 39 ms is

Trial 3's PQ p99 of 39 ms is a real outlier — likely a spot instance CPU steal event or a JVM GC pause, though I haven't isolated the cause yet. Something to nail down in Phase 2 with more trials. Three findings worth stating:

1. In the common case, PQ overhead is

1. In the common case, PQ overhead is essentially invisible. At 300 req/s on Graviton3, X25519MLKEM768 costs the same 2 ms mean handshake as classical X25519. p50 and p95 are the same at ms-resolution. If your SLO is "handshake under 50 ms", PQ vs classical is not what you should be worrying about.

2. The cost shows up in the tail

2. The cost shows up in the tail. Median p99 grew from 5 ms to 8 ms. In one of three trials the PQ p99 was 39 ms — 8x what the median trial saw. This matches what I expected theoretically: KEM operations have more variance than raw X25519, and worst-case scheduling amplifies the difference.

News

PQC-Bench Part 1 — measuring X25519MLKEM768 costs on a $0.04/hour Graviton3

There's a lot of discussion going on about post-quantum TLS.

@spots #dev
Source: Dev.to
See more like this