In a large distributed system with an edge/gateway layer, individual backend services, and per-tenant concerns all in play, how would you design a throttling policy across these layers, and what goes wrong if throttling is only applied at one of them?
answer
- edge = cheap, coarse, abuse shield
- per-service = own dependency bottleneck
- per-tenant = fairness across layers
- gap at one layer = blind spot for that failure mode
- threshold consistency across layers, revisit together
basics
~20 sPut limits at more than one point — at the front door, inside each service, and per customer — instead of just one place. If you only limit at one spot, problems from a different spot slip through and can still overload something downstream.
solid answer
~50 sLayer throttling at the edge/gateway (cheap, coarse, protects the whole system and shields internal services from raw internet traffic and abuse), within each individual service (protects that service's own specific dependencies and capacity, since different services have very different bottlenecks), and per-tenant across the platform (prevents one caller from starving others, which a purely per-service or edge-only limit doesn't address). Each layer catches a different failure mode the others miss: edge-only throttling doesn't protect one overloaded internal dependency from being hammered by traffic that's individually within the edge's global budget; service-only throttling doesn't stop a single tenant from consuming a disproportionate share; and none of them alone gives you graceful, prioritized degradation across the whole request mix during a systemic event. The cost is coordination complexity — thresholds at each layer need to be consistent with each other and revisited together, or one layer's throttle can silently starve traffic another layer thought it was allowing through.
go deeper
Should grasp that limiting traffic in just one place might miss problems that show up somewhere else in the system, in general terms.
Should be able to name at least two layers (e.g., edge and per-service) and give a basic reason each exists.
Should explain the distinct blind spot each layer has and give a concrete scenario of a gap when only one or two layers are present.
Should reason about coordination costs across layers (threshold consistency, revalidation as capacity changes), design prioritized degradation across a full systemic event, and connect layered throttling to overall platform SLA/isolation guarantees.
## Why one layer is never enough A large distributed system rarely has just one place where 'too much load' can hurt it, so a throttling policy confined to a single layer inevitably leaves gaps that match whatever that layer can't see. ## The edge layer The **edge or API-gateway layer** sits in front of everything and is the cheapest place to reject or shape traffic, because a rejection there costs almost nothing — no internal service, database, or downstream dependency is ever touched. It's the right place for coarse, system-wide protection: - absorbing abuse and bot traffic - capping total inbound volume the platform is willing to accept at all - and giving every internal service a baseline shield from raw, unfiltered internet traffic But the edge layer, by construction, doesn't know the internal state of any individual backend service — it can't see that one particular service's database connection pool is close to saturated while every other service behind the gateway is fine, because that's not information the edge has visibility into. ## The per-service layer That gap is exactly what **per-service throttling** closes. Each individual service knows its own actual bottlenecks — a service backed by a slow, hard-to-scale database has a very different real capacity than a stateless service doing simple in-memory computation — and can enforce a threshold tuned to its own specific dependencies. This is essential because a global, edge-level budget that looks perfectly safe in aggregate can still be entirely wrong for one particular downstream service if that service's traffic share happens to spike relative to the rest, or if that one service's own capacity is unusually constrained. Without per-service throttling, an edge-only policy protects the platform's front door but leaves individual internal services exposed to being overwhelmed by traffic that, from the edge's vantage point, looked perfectly acceptable. ## The per-tenant layer Neither edge nor per-service throttling by itself addresses fairness across tenants, which is the third layer's job. A tenant sending disproportionate traffic can stay well within both the edge's global budget and any single service's overall capacity threshold while still consuming far more than its fair share of what's meant to be shared capacity, silently degrading every other tenant riding the same infrastructure. **Per-tenant accounting** — tracking and bounding consumption by caller identity, independent of where in the system that check happens — is the layer that actually protects tenant isolation, and it typically needs to be applied wherever contention across tenants for a shared resource can occur, which might be: - at the edge (a per-API-key quota) - inside a specific service (a per-tenant queue or budget for an especially contended dependency) - or both ## Each layer's blind spot The reason to run all three together rather than picking just one is that each layer's blind spot is a different failure's root cause, and an incident caused by a gap at one layer looks, on the surface, identical to an incident the system is supposed to already be protected against — which makes diagnosis genuinely hard without the layered view. - An edge-only policy misses the case of one specific internal service getting overwhelmed by traffic that's globally fine. - A per-service-only policy (with no edge protection) leaves every internal service individually exposed to raw abuse traffic that a cheap edge check could have absorbed for free. - And a policy with edge and per-service limits but no per-tenant accounting anywhere still lets one tenant degrade everyone else, because neither of the other two layers is looking at identity at all. | Layer | The failure only this layer catches | |---|---| | **The edge or API-gateway layer** | raw abuse traffic that a cheap edge check could have absorbed for free | | **Per-service throttling** | one specific internal service getting overwhelmed by traffic that's globally fine | | **Per-tenant accounting** | one tenant consuming far more than its fair share of what's meant to be shared capacity | ## The real cost is coordination The real cost of a layered design is coordination, not just engineering effort to build multiple mechanisms. Thresholds at different layers have to be kept consistent with each other or one layer can silently defeat another's intent — for instance, if the edge allows a burst that an individual service's own, more conservative threshold immediately rejects, the edge's 'protection' did nothing but add an extra network hop before the same rejection happens anyway, and if the mismatch runs the other way (edge tighter than a service's real headroom), the service is left underutilized despite having spare capacity the edge never let through. Revisiting thresholds also has to happen together: scaling a service's capacity without revisiting its own throttle threshold (or the edge's understanding of that service's share of traffic) leaves the layered system stale relative to what it was actually designed to reflect. ## Prioritized degradation across a systemic event A layered approach also enables prioritized degradation across an entire systemic event in a way no single layer can alone: during a platform-wide spike, the edge can shed clearly abusive or unauthenticated traffic first, individual services can degrade or shed their own lowest-priority endpoints, and per-tenant limits ensure that even legitimate high-priority traffic is spread fairly rather than dominated by whichever tenant happens to be busiest at that exact moment — three coordinated defenses each catching what the other two structurally cannot see.
- Why can't edge-level throttling alone protect an individual backend service from being overwhelmed?The edge only sees aggregate traffic across the whole platform, not the real-time internal state of any one service's specific dependencies, like how close its database connection pool is to saturation. Traffic that looks perfectly safe in aggregate at the edge can still be too much for one particular service if that service's real capacity is unusually constrained or its share of traffic spikes.
- What's a concrete way a mismatch between edge and per-service thresholds can cause a problem, even though both layers are individually 'working correctly'?If the edge allows a burst larger than what an individual service's own, tighter threshold will accept, requests pass the edge only to be rejected immediately inside the service, wasting a network hop and adding latency without actually preventing the rejection. Conversely, if the edge is set more conservatively than a service's real headroom, that service sits underutilized despite having spare capacity the edge never lets reach it.
- Does per-tenant fairness need to be enforced at every layer, or is one layer usually enough?It depends on where contention for a shared resource actually occurs — a per-API-key quota at the edge can catch platform-wide abuse from one tenant, but a specific, heavily contended internal dependency may need its own per-tenant accounting too, since a tenant could pass the edge's aggregate quota comfortably while still monopolizing that one service's scarce resource. In practice it's applied wherever tenants genuinely compete for something limited, not uniformly everywhere by default.
Like airport security with one checkpoint at the terminal entrance, a second check at each individual gate, and a separate no-fly list checked at both — removing any one layer doesn't just weaken protection evenly, it opens a specific hole only that layer was covering.
saying these in an interview costs you the question
- Assumes throttling only needs to exist at one layer (usually just the edge/gateway) to be sufficient
- Doesn't recognize that per-service capacity constraints can differ significantly from what an edge-level budget can see
- Ignores that mismatched thresholds across layers can silently waste capacity or defeat each other
- Treats per-tenant fairness as automatically covered by edge or per-service limits without a specific mechanism for it
- Can't articulate a concrete failure mode unique to skipping one particular layer