Spots

NestJS API Quota Management: meet nestjs-quota

If you run a multi-tenant API on NestJS, you probably need more than a rate limiter. You need to answer "may this request proceed?"

against several limits at once (per user per

against several limits at once (per user per minute, per tenant per day, per tenant per month) and "how much has each tenant actually used?" for billing and dashboards. nest-quota is an open-source package that does both. nest-quota is an open-source package that does both.

It checks and consumes all the relevant quotas

It checks and consumes all the relevant quotas in one atomic operation, so you never end up with a half-charged request. It works across multiple servers using a single Redis Lua script per decision.

It supports idempotent retries, reserve/commit/release for operations whose

It supports idempotent retries, reserve/commit/release for operations whose cost you don't know up front, plan-based dynamic limits, and configurable fail-open or fail-closed behavior. It ships with NestJS decorators, a guard, and an interceptor, and the core has zero runtime dependencies. Redis and NestJS are both optional layers on top. The rest of this post explains why it exists and how it works.

Say you sell an API with three plans

Say you sell an API with three plans. Free gets 10,000 requests a month. Pro gets a million. Enterprise is custom. On top of that, you want to stop any single user from sending more than 100 requests a minute, and stop any single tenant from burning through more than 10,000 in a day. So one incoming request has to be checked against four buckets: The obvious implementation looks like this: This breaks in at least three ways.

Race conditions. Two requests read used = 9999

Race conditions. Two requests read used = 9999 at the same moment, both pass the check, both increment. You just gave away extra usage. With 100 concurrent requests and 10 units left, you can easily let 30 through.

Partial charges. With multiple buckets, you increment the

Partial charges. With multiple buckets, you increment the user/minute bucket, then discover the tenant/day bucket is full, and reject. But the user/minute bucket is already charged for a request that never ran. Now you need rollback logic, and rollback over a network is never truly atomic.

Double charging on retries. A client times out

Double charging on retries. A client times out, retries with the same request, and you charge twice. Your customer notices on their invoice.

On top of that, real APIs have costs

On top of that, real APIs have costs you don't know in advance. An LLM call might use 200 tokens or 8,000. You can only meter that after the handler runs.

That's the gap. Rate limiting answers "is this

That's the gap. Rate limiting answers "is this client too fast?" Quota enforcement and usage metering answer a different set of questions, and they need real accounting semantics. How nest-quota handles it Atomic multi-policy consumption

News

NestJS API Quota Management: meet nestjs-quota

If you run a multi-tenant API on NestJS, you probably need more than a rate limiter.

@spots #dev
Source: Dev.to
See more like this