Spots

Postgres with QUIC: Breaking the Client-Side Connection Bottleneck | Lupyd

Every backend engineer running Postgres at scale eventually learns the same painful lesson: connection pooling does not fix network physics. We deploy PgBouncer or PgCat, configure transaction pooling, lock down backend connections so the database doesn't run out of memory, and celebrate. But if your application or edge services sit tens of milliseconds away from your database cluster, your queries are still choking on a transport bottleneck we rarely talk about: the client-side TCP connection pool.

Recently, I ran an experiment to address this

Recently, I ran an experiment to address this directly: running the Postgres protocol over QUIC (UDP) instead of traditional TCP. We modified PgCat to accept QUIC connections and multiplexed hundreds of virtual database streams over a tiny handful of UDP associations.

The outcome was startling. Under heavy load, our

The outcome was startling. Under heavy load, our QUIC setup pushed 9,718 queries per second at an average latency of 246ms. The identical workload over traditional TCP collapsed into a catastrophic queue pileup of over 40,000 backlogged queries, spiraling to 7,088ms average latency.

Here is the real engineering breakdown of why

Here is the real engineering breakdown of why this happens, the arithmetic of latency asymmetry, where QUIC genuinely changes the rules, and where it cannot save you. The Synchronous Postgres Reality

To understand the problem, you have to look

To understand the problem, you have to look at the PostgreSQL frontend/backend wire protocol (Protocol 3.0). Postgres is fundamentally synchronous at the connection layer.

When a client sends a Query or Execute

When a client sends a Query or Execute message down a Postgres TCP connection, that socket is tied up until the server finishes processing and returns the matching CommandComplete and ReadyForQuery message. You cannot interleave two independent queries from different application threads on the same physical connection without strict sequential serialization.

1 Connection = 1 Active In-Flight Query or

1 Connection = 1 Active In-Flight Query or Transaction. A client socket is completely blocked from the moment bytes leave the network card until the final row batch and ReadyForQuery flag return across the wire.

Because establishing a new Postgres backend process on

Because establishing a new Postgres backend process on the database server involves fork-exec overhead, memory allocation (several megabytes per connection for work_mem, cache metadata, and process state), and catalog locks, we can't let 5,000 application threads open 5,000 direct connections to Postgres. The server would instantly melt from context switching and OOM crashes. So we put connection poolers in the middle. But look closely at where the pooler actually sits. The Latency Asymmetry: 30ms WAN vs 2ms LAN

In modern distributed architectures, your services don't live

In modern distributed architectures, your services don't live in the database rack. You have edge nodes, serverless workers, microservices in different availability zones, or regional clusters in us-east-1 querying a centralized database cluster in us-east-2 or eu-central-1. Let's consider a realistic, standard production deployment: Client to Pooler (WAN / Inter-region): Round-trip time (RTT) is 30ms. Pooler to Postgres (Internal LAN / Same VPC): Round-trip time is <2ms (often sub-millisecond). Query execution time inside the Postgres engine: A fast indexed point lookup or one-liner takes 0.5ms. Each query occupies 1 connection for 30.5ms total waiting on network wire flight <2ms internal LAN latency

The database engine executes the query in 0.5ms

The database engine executes the query in 0.5ms and the internal pooler recycles the connection in 2ms. But the client socket cannot be reused for 30.5ms because the bytes are still flying over the public internet or cross-region fiber! The Math of Client-Side Connection Starvation Here is where basic arithmetic exposes the flaw.

News

Postgres with QUIC: Breaking the Client-Side Connection Bottleneck | Lupyd

Every backend engineer running Postgres at scale eventually learns the same painful lesson: connection pooling does not fix network physics.

@spots
Source: Hacker News
See more like this