throttlekit
·

Scaling to a fleet without changing a line of client code

One global limit across a whole fleet — Tier-1 over the existing RPCs with zero client change, and a Tier-2 lease for the hottest clients. Shipped in throttlekit-server 0.4.0.

A rate limiter that’s correct on one box is the easy version. The moment you run N instances in front of one shared backend, each one limits in isolation and the fleet admits up to what you meant. ThrottleKit’s answer is one oracle, four doors, two tiers, and the headline is that most of it needs no client change at all.

Tier 1: flip a policy block, the fleet coordinates

Four fleet-distributed features now reach any client over the existing gRPC RPCs. You add a block to the server’s config; the client code is byte-for-byte the same:

Policy blockReached overWhat it bounds
federatedCheckone global per-window rate budget across regions
federatedFairEscrowCheckweighted-fair split of one global budget across tenants
fleetBudgetDebitone token budget across every instance
distributedConcurrencyAdmitone in-flight ceiling across the fleet

Each is built on the published core (federate(), distributedTokenBudget, distributedAdaptiveConcurrency, federatedWeightedFairEscrow) over a shared store. That means no wire change, no core release. The decision is always the core’s: the server stays the single oracle, so a Python client and an embedded TypeScript limiter return the same answer.

One caveat worth stating up front: distributedConcurrency needs a unique --node-id per process (it defaults to host#pid). A collision corrupts the fleet aggregate, so identity is mandatory, and a partitioned node self-fences on lease expiry rather than over-admitting.

Tier 2: lease a slice for your hottest clients

A client doing tens of thousands of requests a second doesn’t want a round trip per request. For that, throttlekit-server@0.4.0 adds the additive Fleet door (Fleet.Reserve): the client leases a chunk of a federated: policy’s global budget and spends it locally, round-tripping only to refresh.

The server computes the grant size through the policy’s federation coordinator (again, the one oracle) and the client only spends what it was granted, byte-identically to the core’s leased two-tier L1 (pinned by a golden lease-vector suite every polyglot port replays). It never synthesizes its own denial; it surfaces the server’s. Handing out budget is a poisoning vector, so the Fleet door is loopback-only until you set --fleet-secret.

throttlekit@1.5.0 also closed a subtle correctness gap here: a Tier-2 lease is now anchored to the store clock (leaseWindowed), so a lease holder discards leftover credits exactly at the store’s window boundary. No more node↔store clock skew extending a window past where the store rolls it.

One limit. Any fleet size. And for almost everyone, not a single line of client code to change.