skip to content

An adversary fires 300 near-identical alerts so your response platform hits its API rate limit — how do you stop real containment being dropped?

level: seniorimportance: should knowfreq 34%

answer

  1. 429 means refused, not queued
  2. enrichment burns the shared budget
  3. reserve a lane for containment
  4. dropped action must page, not pass
  5. a cheap flood buys dwell time

basics

~10 s

Give containment actions a reserved quota lane ahead of enrichment, and make a rate-limited containment action halt and page rather than be skipped. Treat the flood as a diversion and hunt that window.

solid answer

~50 s

The failure is that enrichment burns the shared API quota and containment then gets HTTP 429 and is silently dropped. I fix it in three places. First, collapse alerts on the same entity into one case so 300 alerts do not launch 300 runs. Second, split the quota: containment and eradication actions get a reserved, prioritised lane with bounded retries honouring `Retry-After`, and enrichment yields to them. Third, fail loud — a containment action that cannot execute must block the case and page the incident lead, never pass through as if it had run. Then I treat the flood as adversarial until proven otherwise: generating cheap alerts to exhaust the response path is a diversion, so I hunt for what happened quietly during the window, and I check which actions were dropped so those cases can be re-verified.

go deeper

for a junior

Know that an HTTP 429 means the target refused the request because of rate limiting, so the action did not happen, and that a refused containment call needs a human to be told.

for a middle

Explain where the quota goes: enrichment issues far more calls than response, so under load the containment call is the one that collides with the limit unless the lanes are separated.

for a senior

Show the design and the diagnosis together — reserved lanes, bounded retries honouring Retry-After, fail-loud containment, plus re-verifying the actions dropped during the window and hunting what the flood covered.

for a principal

Frame the response platform as a capacity-constrained shared resource an adversary can attack, and decide what share of budget and engineering effort protects the actions whose failure loses a case.

## The shape of the failure A response platform's connectors talk to the same APIs as everything else: the EDR, the identity provider, the artifact repository, the cloud control plane. Those APIs enforce per-tenant rate limits and answer an over-quota call with HTTP 429 (Too Many Requests), usually carrying a `Retry-After` header that says when to try again. A 429 is a refusal — the request was not performed and nothing was queued at the target. Most of a SOC's API budget is spent on *enrichment*, not on response: for every alert, a playbook pulls the asset record, the user's recent sign-ins, the process tree, a handful of reputation lookups. That is dozens of calls per alert against a budget that is shared with the small number of calls that actually change the world. When alert volume spikes, enrichment consumes the quota first and containment — which happens later in the run — collides with the limit. ## Why the flood is cheap for the adversary and expensive for you An adversary does not need to compromise anything to trigger this. Repeatedly touching something a noisy rule watches — a monitored registry path, a flagged binary name, a login pattern that trips an identity detection — produces alerts at essentially zero cost, from a position that does not matter, on hosts that do not matter. Each one launches a playbook that spends real API budget. Three hundred of them are minutes of work. The payoff is asymmetric: the one case that matters arrives inside the flood, its containment action gets a 429, and the case proceeds as if containment happened. Even without exhausting the quota, the flood buries the real case in a queue and buys dwell time. Treat a sudden burst of low-value, near-identical alerts as a possible diversion rather than as a nuisance to be silenced. ## Designing the response path so it survives **Do not launch one run per alert.** Collapse alerts that share an entity — the same host, the same account, the same artifact — into one case, so three hundred alerts drive one investigation and one set of actions. This alone removes most of the load. **Split the quota by consequence.** Route containment and eradication actions through a prioritised lane with a reserved share of the budget, and let enrichment run in a lower lane that yields. Enrichment being late costs you context; containment being dropped costs you the case. **Cap concurrency per playbook class.** A bounded worker pool with a queue is better than unlimited parallelism: unbounded runs sprint into the rate limit together, and every one of them fails. **Retry deliberately, not reflexively.** Honour `Retry-After` and use bounded backoff. Immediate retries deepen the limit for everyone, including the action you care about. Bound the total wait, and when it is exhausted, escalate rather than give up quietly. **Fail loud on the actions that matter.** A rate-limited containment action must set the case to a blocked state and page a human with the specific action that could not be performed. A responder who knows the automation could not isolate the host will do it by hand in a minute; a responder who sees a green run will do nothing. **Keep a path that does not share the quota.** The manual console route, or a separate integration for emergency actions, so an exhausted automated path never means no path. ## Detecting it while it is happening Monitor the automation platform's own execution record as a telemetry source in its own right: the rate of 429 responses per connector, the count of playbook runs queued or dropped, and the time from case creation to first containment action succeeding. A spike in 429s during an incident is an operational alarm and possibly a security signal. Both readings deserve a look, and you cannot distinguish them from the alert stream alone. ## What to do afterwards Two pieces of follow-up work: 1. **Re-verify the affected cases.** Enumerate the actions that returned 429 during the window, check each target's audit trail for whether the action ever landed, and perform the ones that did not. Any case containing a dropped containment action was not contained when it was said to be. 2. **Hunt the window.** Ask what else was happening while the SOC's attention and API budget were consumed. Look at the surfaces the flood did not touch — identity sign-in logs, cloud control-plane audit entries, egress flow records — over exactly that interval. A diversion is only worth running if something is happening behind it, and the flood gives you an unusually precise time box in which to look. ## The distinction to hold onto A broken detection producing three hundred alerts and an adversary generating three hundred alerts look identical for the first ten minutes. The difference is found by evidence — a recent rule change, an unrelated deployment, the spread of affected hosts — not by assumption. Until that evidence exists, the safer default is to assume the volume is intentional and keep the response path protected.

  • How do you tell a deliberate flood from a broken detection producing three hundred alerts?
    By evidence, not instinct. Check for a recent rule or deployment change, whether the alerts cluster on one entity or scatter across unrelated hosts, and whether anything distinct coincides with the burst. A broken rule usually fires on a schedule or a change boundary; a deliberate flood usually keeps touching the same cheap trigger. Assume adversarial until the change is found.
  • Your connector retried 429s for an hour and then succeeded. What is the harm if the action eventually landed?
    Containment happened an hour late, so the case must be dated from the successful call and everything the adversary could do in that hour is in scope. Long silent retries also hide the degradation from responders and keep the platform saturated, so bound the retries and escalate when the bound is hit.
  • Which actions should hold the reserved quota lane?
    Actions that change the state of the estate on an active case: isolate a host, disable an account, quarantine an artifact, block a hash. Enrichment lookups, ticket updates, and chat notifications yield to them. The ranking is by consequence of not running, not by how early the step appears in the playbook.

saying these in an interview costs you the question

  • Raises concurrency, which burns the shared quota faster
  • Turns off automated response during high alert volume
  • Retries rate-limited calls immediately and indefinitely
  • Treats a 429 as the action having been queued
  • Silences the noisy rule without hunting the flood window

context