Four hundred jobs restart together, each asking the store for its own warehouse account — which downstream limits decide whether every one of them starts?
answer
- creation is a write, reads are not
- population, not request rate
- lifetime divided by restart interval
- leftovers hit the cap first
- budget accounts before you need them
basics
~20 sTwo: how fast the warehouse can create a principal, and how many principals it will hold at once. The second is usually hit first, because live accounts are consumers multiplied by how many generations of them are alive simultaneously.
solid answer
~50 sCreating a principal is a **write** to the warehouse's own identity records, often serialised and far slower than the reads that system is sized for, so a synchronised restart turns into a burst on the most expensive path. The limit that actually bites, though, is **population**: many systems cap how many principals may exist. With 400 consumers, accounts that live 8 hours and jobs that restart every 2 hours, four generations overlap — 1600 live accounts, against a cap of 1000. Nothing about today's load explains the failure; it is yesterday's accounts still in place. The fixes are to shorten how long each account lives, reuse one account per consumer for a window instead of per process, stagger restarts, or raise the cap — and to alarm on principal count as a fraction of it long before then.
code
pseudocode · 14 linesconsumers = 400 # jobs that each mint their own account
restartInterval = 2 # hours between restarts of one job
accountLifetime = 8 # hours an account exists before removal
cap = 1000 # principals the downstream system will hold
generationsAlive = accountLifetime / restartInterval # = 4
liveAccounts = consumers * generationsAlive # = 1600
if liveAccounts > cap:
# 1600 > 1000: creation refuses in steady state, with no spike involved
report('over cap by ' + (liveAccounts - cap) + ' accounts')
report('levers: shorten accountLifetime, reuse per consumer, raise cap')
else:
report('headroom: ' + (cap - liveAccounts) + ' accounts')go deeper
Recall that each minted account is a real object in the downstream system, so there is a limit to how many can exist and how fast they appear.
Explain why creation is a write on a slower path than reads, and compute live accounts as consumers times overlapping generations.
Diagnose from the symptom: creation refusals with a healthy data path point at the population cap, so count the principals that exist and check whether removal is keeping up before touching restarts.
Treat account population as a budgeted capacity dimension with an alarm on its fraction of the cap, and decide in advance which systems cannot support the design at fleet scale.
## Two different limits, and they fail differently When a store mints per consumer, two properties of the downstream system decide whether a fleet can start: - **Creation throughput** — how many principals the system will create per unit of time. Creation is a write against the system's own identity records, and on many systems those writes are serialised or coordinated, so the rate is orders of magnitude below what the same system serves for reads. - **Population cap** — how many principals may exist at once. Some systems state one explicitly; others enforce it implicitly through a resource each principal consumes. They produce different symptoms. Exceeding throughput shows up as **slow start-up** that recovers: jobs queue behind the creation path and eventually start. Exceeding the cap shows up as **creation refusals that do not recover**, while queries from already-running jobs stay fast — a shape that misleads people into looking for a network fault, because the warehouse is plainly healthy. ## Why the population, not the rate, usually bites first The count that matters is not how many consumers exist. It is how many accounts are alive at the same time, which is the consumer count multiplied by how many generations overlap: ```pseudocode liveAccounts = consumers * (accountLifetime / restartInterval) ``` With 400 consumers, each account existing for 8 hours, and jobs restarting every 2 hours, four generations are alive at once: 1600 accounts. Against a cap of 1000, the estate is over the limit in steady state, with no spike involved at all. The accounts that consumed the budget were created hours ago by runs that have already finished. This is why the failure is so often misdiagnosed. Everyone looks at the 400 jobs that just restarted. The 1200 accounts already in place are what used up the budget. ## What to check, in order 1. **Count the principals that exist in the system now**, and compare against its cap. If the count is close, nothing about today's traffic is the cause. 2. **Check whether the count is growing monotonically.** If removal is failing or was never arranged, the population only ever rises, and the cap is a deadline you are walking toward at a fixed rate. 3. **Measure the creation path's rate** separately from the data path's. A system can be fast for queries and slow for identity writes at the same time; measuring the wrong one proves nothing. 4. **Compare the account's lifetime against the consumer's restart interval.** That ratio, not the consumer count, is the multiplier. ## The levers, and what each one costs | Lever | Effect on live accounts | What it costs | |---|---|---| | Shorten how long each account exists | Divides the multiplier | More creations, so more pressure on the slow path | | One account per consumer per window, not per process | Removes the multiplier | Several processes share one account, losing per-process attribution | | Stagger or batch restarts | No effect on population | Only relieves the creation burst | | Raise the cap | Buys headroom | Nothing structural changes; the growth rate is unaffected | | Keep a shared credential for that system | Removes the problem | Gives up per-consumer scoping and attribution there | Note the third row. Staggering is the reflex fix and it addresses the *wrong* limit: if the population is the constraint, spreading the same creations over ten minutes changes nothing at all. Diagnose which limit you are against before choosing. ## Designing for it before it happens The account population is a capacity dimension like any other, and it deserves the same treatment: - **Budget it explicitly**: consumers multiplied by overlapping generations, with headroom for a restart of the whole fleet while the previous generation is still alive — which is the worst case and also the ordinary case after a deployment. - **Alarm on it as a fraction of the cap**, not on creation failures. A failure alarm fires when jobs are already not starting; a fraction alarm fires while it is still a planning problem. - **Alarm on failed removals separately**, because a slow leak in removal is invisible until it is a capacity incident. - **Know the cap before you promise the design**. "Can this system create principals?" is the first question; "how many will it hold, and how fast will it make them?" is the second, and skipping it is how a design that worked in a small environment fails on the day it meets the real fleet.
- Creation is refusing but queries from running jobs are fast. What does that shape tell you?That the data path is healthy and the identity path is not, which points at the population cap rather than at load, a network fault or an outage. Count the principals that exist and compare against the system's limit before touching anything else; if the count is near the cap, the cause is accounts created hours ago, not the jobs that just restarted.
- Why is staggering the fleet's restarts often the wrong fix here?Because it addresses creation throughput, and the constraint is usually population. Spreading the same 400 creations over ten minutes leaves exactly the same number of accounts alive at the end of it. Staggering helps only when the symptom is slow start-up that eventually succeeds.
A car park does not care how fast cars arrive; it cares how long each one stays. Halve the stay and you double the capacity without laying a single extra bay.
saying these in an interview costs you the question
- Sizes the downstream by request rate, not live account count
- Assumes creating a principal costs what reading a row costs
- Reads a creation refusal as the warehouse being down
- Ignores accounts left behind by earlier runs
- Raises the cap without changing lifetime or reuse