One tenant's export monopolises a shared analytics query pool — which denial-of-service controls answer this threat?
answer
- the offender is a legitimate paying user
- counting requests measures the wrong thing
- bound what one tenant can hold
- stop sharing the exhausted resource
- concurrency slots, cost ceilings, separate lanes
basics
~20 sAuthentication cannot help: the offender is a legitimate tenant. Answer with cost-based quotas and isolation - per-tenant concurrency slots, query timeouts and result ceilings, separate lanes for batch and interactive work, and shedding by priority.
solid answer
~50 sThe adversary is authenticated, paying, and often not malicious at all — a badly written nested export that runs for minutes. That rules out the usual identity controls and points at three families. **Quotas that measure cost, not count**: concurrency slots per tenant, wall-clock timeouts, ceilings on scanned bytes, returned rows and nesting depth, plus a budget over a rolling window rather than an instantaneous rate. **Isolation**: separate queues or worker pools per workload class so a long export cannot occupy the workers serving interactive dashboards, with the noisy tenant queueing in its own lane rather than the shared one. **Shedding by priority**: when the pool saturates, defer exports before interactive queries, and reject fast rather than buffering. Under all three sits per-tenant measurement - you cannot enforce a budget you do not meter. Write the threat against the shared resource, not the endpoint: the pool is what is exhausted.
go deeper
Be ready to explain that one customer of a shared service can starve the others, and that limits on how much work each customer may consume are what stops it.
Explain why counting requests is the wrong unit, and name concrete cost-based limits: concurrency slots, timeouts, rows or bytes scanned, nesting depth, and a budget over a rolling window.
Show you would isolate the lanes as well as bound the tenant — separate queues for batch and interactive, fair scheduling, shedding by priority, fast rejection at the queue bound — and that you measure per tenant before enforcing.
Own the tenancy model: which classes of work share infrastructure, what limits are published to customers as part of the contract, and when a large tenant graduates to dedicated capacity instead of tighter throttles.
## Why this threat is different Most denial-of-service findings assume an outsider. This one does not. The actor is an authenticated tenant of a shared multi-tenant service, exercising a feature they are entitled to use, and the victims are the other tenants on the same infrastructure. Three consequences follow immediately, and a senior answer names all three. First, **identity controls contribute nothing**. Stronger authentication, mutual TLS, better session handling — none of it changes the outcome, because the caller is exactly who they claim to be and is authorised for the operation. The control has to shape consumption, not access. Second, **the offender is frequently honest**. The usual trigger is a scheduled export with a deeply nested grouping that a customer built themselves, or a report that got slower as their data grew. So the mitigation must be one you are willing to apply to a good customer: it has to degrade them fairly and predictably, with a clear error, not ban them. Third, **the threat belongs to the shared resource, not the endpoint**. Write it that way: *an authenticated tenant submits a query whose cost is unbounded relative to the request, occupying shared worker slots for minutes and denying interactive queries to every other tenant.* Naming the pool is what makes the mitigation obvious; naming the endpoint sends you off writing per-endpoint patches. ## Cost-based quotas beat count-based limits A request-per-second cap bounds arrivals. It says nothing about what an arrival costs, and in an analytics API the variance between the cheapest and most expensive permitted query is enormous. The limits that bite are the ones denominated in the resource actually being exhausted: - **Concurrency slots per tenant** — the single most effective control, because it directly bounds how much of the pool one tenant can hold at any instant. - **Wall-clock timeouts** on every query, with a lower ceiling for interactive paths than for batch ones. - **Work ceilings**: bytes or partitions scanned, rows returned, result size, join and nesting depth, expansion factor. Reject at plan time where you can — refusing a query before it runs is far cheaper than killing it halfway. - **A budget over a window**: units of work per tenant per hour, so a tenant who is within their concurrency limit but running continuously still hits a ceiling. - **Admission control**: when the pool is above a threshold, refuse new expensive work rather than admitting it and thrashing. ## Isolation: stop sharing the thing being exhausted Quotas bound a tenant; isolation removes the coupling. Practical forms, cheapest first: separate queues per workload class so exports never take a worker reserved for interactive queries; a bulkhead of dedicated workers per tenant tier so the paid tier cannot be starved by the free one; a hard cap on the fraction of the pool any single tenant may occupy; and, at the far end, dedicated infrastructure for the largest tenants. Fair scheduling sits between quotas and isolation — service the queue by tenant in rotation, or by a weighted budget, so a tenant with a thousand queued exports advances at the same rate as a tenant with one. ## Degrade deliberately When the pool saturates anyway, what fails should be a design decision, not luck. Shed by priority: kill or defer batch exports before interactive dashboards, and return a clear, retryable error with a suggested time rather than a timeout. Reject fast when the queue is at its bound — an unbounded queue converts a compute shortage into a memory exhaustion and moves the outage rather than containing it. ## Denial of service with no attacker in the same system The mirror case is worth modeling alongside: the system exhausting itself. Every client retrying the same failed query at the same moment with no jitter; a synchronised refresh that has every device wake at the top of the hour; a cache expiring for all keys at once so every miss stampedes the same backend. The trigger is not an adversary — it is the system's own traffic, amplified by a shared clock or a shared failure. The controls are their own family: retry with exponential backoff **and jitter**, request coalescing so a thousand identical misses become one backend call, circuit breakers that stop the retries during an outage, and staggered schedules instead of round numbers. ## What to expect to be pushed on An interviewer will usually challenge the easy answer. If you say rate limiting, they will ask why one permitted query holding a worker for four minutes is bounded by it. If you say autoscaling, they will ask what happens in the minutes before capacity arrives, and who pays for it. If you say block the tenant, they will ask how you explain that to a paying customer whose report simply got slower. The defensible position is layered: measure per tenant, bound by cost, isolate the lanes, shed by priority, and make the limits documented and visible so a customer can design within them.
- Per-tenant rate limiting is already in place. Why does the threat survive?A rate limit caps how many requests arrive, not what any one costs, and one permitted query can hold a worker for minutes. You need limits denominated in the exhausted resource — concurrency slots, timeouts, scanned-byte and result ceilings, nesting depth — plus isolation so an over-budget tenant queues in its own lane instead of the shared one.
- How does this differ from an anonymous attacker exhausting the same pool?The identity is known and legitimate, so cutting them off is a commercial decision rather than a security one, and the control must degrade them fairly and explain itself. The trigger is often accidental, which means the mitigation has to hold for honest users too, and the limits must be documented so customers can design within them.
- What does a denial-of-service threat with no attacker look like in this same system?Self-inflicted amplification: every client retrying a failed query at the same instant without jitter, or a synchronised refresh that wakes every consumer at once, so the service's own traffic exhausts it. Model it as a D against the shared resource and answer with backoff plus jitter, request coalescing, circuit breakers and staggered schedules.
- Where should the enforcement point live?At admission to the shared resource, not at the API edge alone. The edge does not know what a query will cost; the planner or scheduler does. Enforce cost ceilings at plan time so an over-budget query is refused before it runs, and hold the concurrency slots at the pool that is actually contended.
saying these in an interview costs you the question
- Says stronger authentication or mutual TLS fixes it
- Relies on request-rate limits for cost-heavy queries
- Assumes autoscaling makes exhaustion impossible
- Treats it purely as a capacity-planning problem
- Ignores that the noisy tenant may be entirely honest
- Adds an unbounded queue and calls it backpressure