A platform runs many ingest pipelines in one process; how do you decide how many execution contexts it exposes and which workloads may use each?
answer
- start from blast radius
- shared inventory versus per-workload contexts
- isolation costs threads and memory
- split on evidence, not prediction
- per-context queue depth on a dashboard
basics
~20 sDecide by blast radius rather than convenience. A small shared inventory keeps the process's total worker count predictable and understandable; a dedicated context per workload buys isolation, paid for in threads, idle memory and harder global reasoning.
solid answer
~50 sThe two ends are clear, and the judgement is where to sit between them. A small shared inventory — one compute context, one waiting context, a serial worker per resource that needs one — keeps the total number of workers on the host a number you can state, but it couples every team to every other: one workload's slow calls consume the capacity the rest depend on. A context per workload gives isolation, and pays for it with more workers than the host has processors, stacks that sit idle, and a process nobody can reason about as a whole. I would start shared, then split out a workload when there is evidence of coupling — an incident where one workload's latency reached another — or when two workloads have genuinely different latency goals. Then I would make the rule enforceable rather than documented, and put worker count and queue depth per context on a dashboard.
go deeper
Know that a process does not get its own worker pool per feature by default; execution contexts are shared capacity somebody decided about.
Explain what workloads sharing a context do to each other: they queue behind one another and compete for the same finite set of workers.
Bring evidence to the argument — per-context queue depth and worker counts, plus the incident where one workload's slowness reached another through a shared context.
Own the trade-off and its enforcement: how many contexts exist, who may add one, what happens to a workload fitting none of the kinds, and what the whole standard costs in threads.
## The choice, stated honestly Execution contexts are shared capacity inside one process. Deciding how many exist is deciding how much isolation the workloads have from each other, and isolation is not free: it is bought with threads, memory and complexity. The two ends: | | Small shared inventory | A context per workload | |---|---|---| | Total workers | predictable, a number you can state | grows with the number of teams | | Blast radius | one workload's slowness becomes everyone's queueing time | contained to the workload that caused it | | Reasoning | one picture of the process | many small pictures, no global one | | Idle cost | workers are shared and stay busy | stacks parked in contexts that are quiet | | Failure mode | shared capacity exhausted by one tenant | oversubscription; more runnable workers than processors | Neither end is the answer. The answer is a default plus a documented route for exceptions. ## What should decide it 1. **Demonstrated coupling.** Has one workload's slowness already reached another through a shared context? That is the strongest argument for splitting, and it is evidence rather than prediction. 2. **Different latency objectives.** A workload serving an interactive request and a workload doing background repair should not queue behind each other, even when both are healthy. 3. **Trust in the dependency.** Work that calls something with a history of slowing down deserves its own bounded context, so that its bad days are its own. 4. **Count of workloads.** Ten workloads with a context each is a different process from two hundred with a context each; at some scale the shared inventory stops being a preference and becomes arithmetic. 5. **Ability to observe.** A context you cannot see the queue depth of should not be created, because a split you cannot measure has only moved the problem. ## Making the standard hold A standard that lives in a document is a standard that decays. What makes it hold: - **Make the right choice the easy one.** Hand teams a small factory that returns the context for a named workload kind, so placing a stage is a call rather than a judgement each engineer re-makes. - **Publish per-context metrics as first-class.** Worker count, queue depth and time spent queued, labelled by context. A misplacement becomes visible as a shape on a chart instead of an argument in review. - **Make adding a context a reviewed change with a named owner.** Contexts multiply quietly otherwise, and a process with forty of them has no inventory at all. - **Bound every context that can grow.** A ceiling per context is how the isolation claim becomes true rather than aspirational. ## The exceptions the standard must name Every real standard meets a workload that fits none of the kinds — one that computes substantially and then makes a slow call inside the same step. If the standard is silent, that workload lands wherever is nearest, and it is usually the shared waiting context, where its compute work quietly competes with everyone's calls. The standard should name the route instead: split the step so each half runs on its matching context, and where the split is genuinely impossible, give that workload a context of its own with its own bound so the mismatch stays contained. A second exception worth naming ahead of time is the workload that needs one-at-a-time execution. Whether each such resource gets its own serial worker or they share one is exactly the same isolation trade-off in miniature, and it is better settled once in the standard than argued per feature. ## Knowing whether the decision was right The measure is not thread count; it is whether incidents stay where they started. Two signals are worth tracking over a few months: how often an incident in one workload produced latency in an unrelated one, and how often a team had to request a new context because the existing kinds did not fit their work. The first rising says the inventory is too shared; the second rising says the kinds are wrong, not that more contexts are needed. That distinction is the part of this decision a lead owns, because nobody looking at a single pipeline can see it.
- What makes a context standard hold in practice rather than on paper?Make the correct choice the path of least effort. Give teams a small factory returning the context for a named workload kind, so placement is a call rather than a fresh judgement; publish worker count and queue depth per context so misplacement shows up as a shape on a chart; and make adding a new context a reviewed change with an owner. A rule nobody can see being broken will be broken.
- A workload computes heavily and also makes one slow call in the same step. What does the standard say?It should name a route rather than leave a gap. Split the step so the computing half and the waiting half each land on their matching context; where that split is genuinely impossible, give the workload its own context with its own bound so the mismatch cannot reach anyone else. A standard listing only the clean cases pushes every awkward workload onto whichever context is nearest.
- Why not simply give every team its own contexts and avoid the argument?Because isolation is paid for in threads. Each context carries workers, stacks and scheduling presence whether or not it is busy, and a process with dozens of them has more runnable workers than processors and no picture anyone can reason about. Per-workload contexts are the right answer for the few workloads that have earned one with evidence, not the default for all of them.
saying these in an interview costs you the question
- Give every team its own contexts; isolation is always worth it
- One context for everything, because simplicity wins
- Writing the rule down is enough to make teams follow it
- Total worker count does not matter while most workers are idle
- Which context a stage uses is a per-team implementation detail