skip to content

A tenant on a shared in-memory store deletes its entire prefix at peak and every other application times out - why, and how should it have been done?

level: seniorimportance: should knowfreq 45%

answer

  1. someone else's incident, not yours
  2. capacity is shared while it runs
  3. bounded batches with pauses between
  4. announce first, measure caller-side

basics

~20 s

The removal consumes the shared instance's execution capacity while it runs, so every other tenant's calls queue behind it. The safe form is bounded batches with pauses between them, run off peak, with the co-tenants told beforehand.

solid answer

~50 s

Removing a whole prefix is either a very large number of operations or one operation that finds and removes them all, and either way the work is done by the one shared server. On a store whose server executes one operation at a time, every other tenant's call waits behind it; on a multi-threaded store the effect is softer but capacity is still finite and still shared, so the tightest-timeout caller fails first. Locating the entries is its own hazard: a single operation that returns the whole keyspace builds one enormous reply and is itself an outage. The safe procedure is to traverse incrementally in bounded batches, remove each batch, pause between batches, run it outside peak, and watch caller-observed latency rather than server-side execution times. On a shared tier this is a change with someone else's blast radius, so it is announced, not just scheduled.

go deeper

for a junior

Take away the rule rather than the procedure: on an instance other applications use, a large cleanup is not a private action, and it belongs outside peak hours.

for a middle

Explain where the time goes - one shared pool of execution capacity, a single enormous reply if the keys are located the wrong way - and describe batching with pauses as the fix.

for a senior

Run it safely and prove it: announce the window, traverse incrementally, size batches against caller-observed latency, and know which of the co-tenants' timeouts fails first.

for a principal

Treat the event as a missing bound rather than a careless engineer. Decide where large removals are gated for everyone and which tenants have earned their own instance.

## Why a tenant's own cleanup is everyone's incident The tenant did nothing conceptually wrong: those were its entries, and removing them is the correct end of their life. What makes this an incident is that **the work is done by a server nobody divided up**. A shared instance has one pool of execution capacity, and a bulk removal is a large amount of work submitted into it in a short time. The neighbours did not send that work, cannot see it, and have no queue of their own to be served from. Three separate costs are in play, and a good answer separates them: 1. **Finding the entries.** Unless the tenant already knows every key, the keys have to be located. A single operation that returns the whole keyspace builds one reply proportional to the number of entries, holds the server while it does so, and then ships it - on a busy instance that alone is the outage. 2. **Removing them.** Each removal is cheap; a few million of them are not. The instance is now serving one tenant's cleanup and everyone else's traffic from the same capacity. 3. **Giving the memory back.** Freed entries return memory to the process's allocator, and the resident footprint the operating system sees does not necessarily fall with the stored-data size. A co-tenant watching the wrong number concludes nothing happened. ## Where the variation is The first instinct - 'the server runs one operation at a time, so everything blocks' - is true of one design, not of the class: - On **single-execution-thread** servers, one long operation is a stop-the-world event for every caller of that instance, however unrelated their keys. - On **multi-threaded** servers that reach per-operation atomicity by locking the entry, unrelated callers can make progress, but the cleanup still consumes cores, memory bandwidth and connection capacity. Bulk work saturates a multi-threaded server perfectly well; it just degrades more gradually. - If the instance has a **copy** following it, every removal is propagated. A copy that was keeping up now has a long stream to apply and falls behind, and callers reading from it see entries the primary no longer holds. On stores that acknowledge a write only after a copy holds it, the primary's own writes slow instead of the copy drifting. ## The safe procedure on a live shared tier 1. **Agree the window.** On a shared instance this is a change to a system other teams depend on. Tell the co-tenants what will run, when, and what to watch. The point is not politeness - it is that they, not you, will see the symptom first. 2. **Locate incrementally.** Walk the keyspace with a cursor in bounded batches rather than asking for all of it at once, so the server's work per call stays small and predictable. 3. **Remove in batches with pauses.** A batch size and a pause between batches are the two knobs that turn an outage into a background task. Start small, watch, then increase. 4. **Measure from the caller.** Watch latency as the co-tenants experience it, including time spent waiting for a connection. The server's record of operations that exceeded a time threshold records execution only, so it can look untroubled while callers are timing out. 5. **Stop on a threshold, not on a feeling.** Decide beforehand what caller-observed latency ends the run, and be willing to finish tomorrow. ## What each approach costs | Approach | Effect on co-tenants | When it is defensible | |---|---|---| | One operation that returns the whole keyspace, then remove everything | Worst: a long hold plus an enormous reply, at peak | Never on a live shared instance | | Remove the whole prefix in one pattern-matched operation | Long single unit of work the neighbours queue behind | Small keyspaces, off peak, on an instance you own alone | | Incremental traversal, batched removals, pauses | Bounded and tunable; can be stopped at any point | The default on any shared tier | | Restart or empty the instance | Total: every tenant loses everything at once | Only with every occupant's agreement | ## The platform's part of the failure It is worth saying out loud in an interview that the tenant is not the only party at fault. **This class of store generally offers no per-tenant rate limit**, so nothing prevented an ordinary cleanup from becoming a shared outage. What a platform can do is route tenants through a client library or proxy it controls and gate large removals there, publish a cleanup procedure with batch sizes and a window, and move any tenant whose cleanups are routine onto its own instance - the one bound that does not depend on everyone remembering the procedure. ## Afterwards Record the caller-observed latency during the run and the batch size that produced it, because the next cleanup is sized from that number rather than from a guess. And check the resident footprint against the stored-data size before anyone concludes the memory was not actually released.

  • The instance has a copy following it - what does the bulk removal do there?
    Every removal is propagated, so a copy that was keeping up acquires a long stream to apply and falls behind. Callers reading from it see entries the primary has already dropped, and a failover during the backlog can lose the tail of the stream. On stores that acknowledge a write only once a copy holds it, the primary's own writes slow instead of the copy drifting.
  • How do you stop the next tenant from doing this without asking?
    On a store with no per-tenant rate limit - which is most of this class - you cannot stop it at the server. What a platform can do is put large removals behind a client library or proxy it controls, publish a cleanup procedure with batch sizes and a window, and move any tenant whose cleanups are routine onto its own instance.
  • Stored-data size fell during the cleanup but the resident footprint did not - what does that gap mean?
    The server released the entries and the process kept the memory. Memory returned to the allocator is not necessarily returned to the operating system, so the two numbers diverge after any large removal. It is not evidence the cleanup failed, and it is a reason never to judge a removal by the footprint alone.

saying these in an interview costs you the question

  • Assumes removing one's own keys can only affect one's own application
  • Reaches for a single operation that returns the whole keyspace to find the keys
  • Judges the impact from the server's execution-time record alone
  • Thinks a multi-threaded server makes a bulk removal free for the neighbours
  • Runs a large cleanup at peak because the entries are only a cache
  • Restarts the shared instance to clear the entries quickly