skip to content

Dozens of jobs across a platform write into the same shared operational stores. What house rules stop any single run from flattening one?

level: principalimportance: nice to knowfreq 36%

answer

  1. budget the destination, not the job
  2. one store, many jobs
  3. measure the knee, do not guess
  4. specify the exhausted case
  5. someone must own the headroom

basics

~20 s

Make the write budget a property of the destination, not of each job: a measured concurrency and rate allowance per store, divided among the jobs that write to it, with batching required, writer concurrency decoupled from the piece count, and a defined behaviour when the allowance runs out.

solid answer

~40 s

Per-job tuning fails here because the resource being protected is shared: every team can be individually reasonable and still add up to an outage. So the budget is defined per destination — a simultaneous-writer count and a rows-per-second ceiling the store was measured to sustain — and allocated to the jobs that write into it. The rules that make it enforceable are dull: writers batched, writer concurrency set by policy rather than inherited from the piece count, and a stated behaviour when the allowance is exhausted (slow down, fail fast, or divert to files for a later bulk load). The hard parts are not technical: discovering the real ceiling by measurement, naming an owner for the store's headroom, and paying the migration cost of retrofitting jobs that already run.

go deeper

for a junior

The takeaway to recall is that a shared store is a limited resource with other users, so how fast a job may write it is a platform decision rather than a per-job setting.

for a middle

Be able to explain why writer concurrency must be decoupled from the piece count, and why a rate ceiling has to apply to the whole run rather than to each writer.

for a senior

Show that you would measure the store's knee rather than trust a published figure, and that you would specify what a job does when its allowance runs out.

for a principal

Own the trade-off across shapes — static caps, a central limiter, a buffer, bulk loads — with the coordination cost, the wasted headroom, the migration bill and the exemption path all named.

## The budget belongs to the destination The single move that makes this tractable is to stop attaching write limits to jobs and attach them to destinations. A job's own view is local — its rows, its deadline, its piece count — and dozens of locally reasonable jobs can still saturate one store. A store, by contrast, has one finite thing to spend: concurrent writers it can serve without its latency degrading for everyone else. Budget the thing that is scarce. Concretely, a destination carries two numbers and one rule: - a **simultaneous-writer allowance**, the number of in-flight writers the store tolerates; - a **sustained rate ceiling**, rows or requests per second across all jobs; - an **allocation**, saying how those are divided among the jobs that write there. ## What the house rule actually contains 1. **Writes are batched.** Per-request overhead is paid once per batch, and a per-row writer is treated as a defect rather than a tuning opportunity. 2. **Writer concurrency is set by policy, never inherited.** The piece count was sized for compute — bytes per piece against available worker threads — so letting it decide how many writers hit a store is an accident waiting to be repeated. A job with 20,000 pieces and an allowance of sixteen writers writes them in turn. 3. **A rate ceiling applies to the run, not to the writer.** Otherwise the ceiling scales with concurrency, which defeats it. 4. **The exhausted case is specified.** Slow down and take longer, fail fast and be retried later, or divert to files for a bulk load are three different contracts with three different consumers; leaving it unstated means every job resolves it differently under load. 5. **Bulk paths are preferred where the store has one.** A job that writes files and hands the store one load takes no writer allowance at all. ## Discovering the number instead of guessing it A vendor's headline throughput figure describes an idle store doing the simplest possible write. The number this policy needs is different: the concurrency at which **latency for the store's existing users** begins to degrade. That is found by measurement — ramp writers against a representative store while watching read latency, and take the knee, with margin. It also drifts, because the store's own workload changes, so the number is re-measured rather than set once. ## The policy shapes, and what each costs | shape | how it holds | what it costs | |---|---|---| | static per-job caps | each job configured with its own writer limit | wastes headroom when only one job runs; drifts as jobs are added | | central rate limiter | writers take permits from one service | a coordination point on every write, and a new dependency that can fail | | queue in front of the store | jobs write to a buffer, a consumer drains it at a fixed rate | absorbs bursts well; adds a component, a delay, and its own failure modes | | files plus bulk load | jobs never write rows to the store at all | needs a staging location and a second step, and not every store has a load path | Most platforms end up with static caps as the floor and one of the other three for the destinations that actually matter, rather than one mechanism everywhere. ## The organisational half, which is the real question - **who owns the store's headroom** — if it is nobody, the budget is fiction the first time a deadline is at stake; - **how a new job learns its allowance** — discoverable defaults beat a document, because the default is what ships; - **what retrofitting costs** — hundreds of existing jobs written against no limit, each needing a change and a re-test, is the reason this decision is usually made after the first incident rather than before it; - **what is deliberately exempt** — a nightly job with the store to itself does not need the same ceiling as one running at peak, and a policy with no exemption path is routed around. One engine-level caution belongs in the answer: runtimes differ in whether writer concurrency can be capped independently of the run's division at all. Where it can, this is configuration; where it cannot, honouring the allowance means restructuring the job — merging pieces before the write, or routing through a buffer — and that difference is worth confirming before a policy promises a number the jobs cannot actually hold. The subjects next door are not this one: scheduling those jobs against each other is `workflow scheduling`, and what the destination sees if a failed attempt is retried is `The Visible Effect`.

  • What should happen when a job's write allowance is exhausted?
    It has to be a stated choice, because the three options are different contracts: slow down and miss a deadline, fail fast and be retried in a later window, or divert to files for a bulk load. Unstated, each job improvises under load, which is exactly when consistency matters most.
  • What does a central rate limiter cost compared with static per-job caps?
    It buys accurate global enforcement and pays with a coordination point every writer consults, plus a new dependency whose own failure has to be designed for. Static caps need no coordination but waste headroom when only one job runs and drift out of date as jobs are added.
  • Why is a vendor throughput figure the wrong basis for the ceiling?
    It describes an idle store doing the simplest write, while the policy needs the point at which the store's existing users start to see worse latency. That knee is found by ramping writers against a representative store, taken with margin, and re-measured as the store's own workload changes.

saying these in an interview costs you the question

  • Sets a write limit per job rather than per destination.
  • Takes a headline throughput figure as the safe sustained rate.
  • Leaves each team to find the store's ceiling by breaking it.
  • Assumes a central rate limiter is free of coordination cost.
  • Defines no behaviour for the case where the allowance is exhausted.
  • Writes a policy with no exemption path and expects it to be followed.